A caching discount reaches my bill only when the task actually reuses context. Taking the biggest percentage in an announcement and applying it to all consumption produces excellent savings on paper. The invoice tends to have a less generous interpretation of the arithmetic.
On September 1, 2026, Anthropic announced Fable 5.1 and Mythos 5.1. Fable kept input at $10 and output at $50 per million tokens. Cache reads fell to $0.25, a 75% reduction. The company estimated savings of 25% to 45% on its workloads. Ordinary input still has a different price, and the composition of each execution determines how much the discount matters.
On September 3, I asked whether simple tasks were going to the cheapest model. I was questioning how work was assigned: I wanted to know whether we were paying for capability where it made a difference. Fable's announcement adds a concrete variable. Two similarly sized assignments can cost differently if one repeatedly returns to the same documents while the other receives new material.
For an estimate, I want reused reads separated from new content and generated answers. A long investigation might carry stable instructions across several calls. A short request might involve little repetition. Applying the same percentage to both conceals the question that should guide routing: what consumption is required to reach the expected result?
In June, I had already objected to passing every project's context between agents. Cheap caching has not changed my opinion. An irrelevant discussion still takes up space and blurs responsibilities when reading it costs less. I will keep common instructions separate from request-specific material. That also makes it easier to inspect what reached the agent when something goes wrong.
An additional execution needed to repair the result belongs in the bill. If a cheap choice forces another model to repeat the investigation, I need to include that work. A more expensive option likewise deserves credit when it finishes with less intervention. The announcement cannot decide which of those outcomes will occur.
Fable 5.1 goes on my shortlist for comparing tasks with repeated context. My routing decision will require cost through to an accepted deliverable, with token categories separated. The 25% to 45% remains Anthropic's estimate. Before it becomes savings in my work, the discount has to appear in an execution whose result I can review.