Anthropic’s pricing page for Fable 5.1 has a number that looks great in a slide deck: cache reads dropped from $1 to $0.25 per million tokens, a 75% cut. The headline claim is up to 45% savings on heavily agentic workloads. Then Artificial Analysis ran its own numbers and reported the opposite for max-effort tasks: 20% more expensive per task than Fable 5, because Fable 5.1 generates roughly 1.7x as many output tokens to get there.

Both numbers are correct. They’re just describing different workloads. I spent an afternoon running our own agent loop’s token logs through both scenarios to figure out which story applies to us, and the exercise is worth walking through because the same math applies to whatever agent you’re running.

Why the same price cut can save you money or cost you more

Input and output tokens on Fable 5.1 are priced at $10/M and $50/M — same as Fable 5, no change there. The entire pricing story lives in two places: cache reads dropped 4x, and the model tends to think longer, which means more output tokens per task.

For an agent loop, those two effects pull in opposite directions:

  • Cache reads scale with turns. Every turn in a long agentic run re-sends your system prompt, tool definitions, and conversation history. If most of that is cached, cutting the cache-read price by 75% multiplies directly against your highest-volume cost line.
  • Output tokens scale with reasoning depth. A model that “thinks more” per turn burns through the $50/M output rate faster, and at 1.7x the output volume, that’s not a rounding error — it can outrun the cache savings entirely.

Which effect dominates depends entirely on your loop’s shape: how many turns, how much of each prompt is cacheable, and how verbose the model gets at whatever effort level you’re running.

Running the numbers on our own loop

I pulled logs from one of our internal agents — a code-review bot that runs roughly 15 tool-calling turns per PR, with a system prompt plus tool schema of about 8,000 tokens that’s identical on every turn (a textbook cache candidate).

Before (Fable 5, standard effort), per PR review:

  • Cache reads: 14 turns × 8,000 tokens × $1/M = $0.112
  • Output: ~15 turns × 600 tokens avg × $50/M = $0.45
  • Total cache + output: ~$0.56

After (Fable 5.1, standard effort), per PR review:

  • Cache reads: 14 turns × 8,000 tokens × $0.25/M = $0.028
  • Output: 15 turns × 600 tokens avg (standard effort doesn’t trigger the 1.7x blowup Artificial Analysis measured at max effort) × $50/M = $0.45
  • Total cache + output: ~$0.48

That’s a 14% reduction for us — real, but nowhere near the “up to 45%” headline, because our per-turn output volume didn’t change at standard effort. The 45% figure and the 20%-worse figure are both edge cases: one assumes cache-read cost dominates your bill (true for long loops with small, cheap outputs), the other assumes you’re running at max effort where the model’s longer chains of thought eat the savings (true for hard reasoning tasks with short prompts).

Our code-review bot sits in the middle: moderate cache reuse, moderate output. The lesson isn’t “you’ll save 14% too” — it’s that you have to run your own logs through this, because the vendor benchmark and your workload are very likely different shapes.

What actually matters when you re-profile a model swap

A few things I’d tell any tech lead about to flip a model version in production:

Separate your turns by effort level before you average anything. If you route some calls at standard effort and escalate to max effort for hard cases — which is the tiered-routing pattern most of us run now — a blended average across effort levels will hide the max-effort cost spike entirely. Pull logs per effort tier, not per model.

Cache hit rate matters more than cache price. A 75% price cut on cache reads does nothing if your system prompt changes on every turn, invalidating the cache. Check your actual cache-hit percentage before assuming the discount applies to you. Ours sits above 90% because the tool schema and system prompt are static; a bot with a dynamic system prompt built per-request would see almost none of this benefit.

Output token growth compounds with self-correction loops. If your agent retries or self-corrects — ours does, twice on average per PR when static analysis flags something — the 1.7x output multiplier at max effort applies to every retry, not just the first pass. That’s where “20% more expensive” can turn into 35% or 40% for retry-heavy workloads.

Where we landed

We kept Fable 5.1 at standard effort for the code-review bot — the 14% saving is real and the model’s suggestions got noticeably better, which matters more to us than the last few cents per PR. We’re holding off on escalating any of our max-effort research agents to Fable 5.1 until I can run the same log analysis on those loops, because those are exactly the workloads where Artificial Analysis’s 20%-worse number is more likely to be our number too.

The broader point: vendor pricing announcements describe an average across workloads nobody actually runs. Your bill is the average of your workload, not theirs. Pull your logs before you trust either headline.

Export for reading

Comments