Spots

Fable 5.1 costs twice as much as Opus 5, until your cache gets big enough

Claude Fable 5.1 shipped on September 1 with the same $10 input / $50 output per million tokens as Fable 5. One number moved: cache reads dropped from $1 to $0.25 per million.

On the price sheet it still costs twice

On the price sheet it still costs twice what Opus 5 does ($5 / $25). But look at the cache-read column and the order flips. Fable 5.1 charges $0.25 there, Opus 5 charges $0.50. So which one is cheaper depends on how much of each request is served from cache, and I wanted the actual crossover instead of a vibe. A request bill has three parts

Cached tokens times the cache-read price, fresh input

Cached tokens times the cache-read price, fresh input times the input price, output times the output price. Fable 5.1 only changed the first term, so the saving over Fable 5 is exactly as big as that term's share of your bill. Two workloads I priced out from the table on my site:

In the agent loop, cache reads are about

In the agent loop, cache reads are about 60% of the Fable 5 bill, so the price cut takes 44% off each request. That lines up with the "around 45% for heavy agent use" figure in the launch coverage.

The chat case barely moves. Output dominates, and

The chat case barely moves. Output dominates, and Opus 5 stays roughly half the price. Switching there is just paying more. Fable 5.1 pays double for fresh input and output, and half for cache reads. Set the two bills equal:

R is cached tokens per request, U is

R is cached tokens per request, U is fresh input, O is output. Plug in the agent numbers above and the line sits at 140K cached tokens. Both models cost $0.105 a request there.

Past that point the gap keeps widening. At

Past that point the gap keeps widening. At 200K cached, Fable 5.1 is about 11% cheaper. Double the cache again and it's more than a quarter cheaper.

Output length is what pushes the line out

Output length is what pushes the line out. Every extra thousand output tokens needs another hundred thousand cached tokens before Fable 5.1 catches up. The part I got wrong the first time

My first pass stopped there and the conclusion

My first pass stopped there and the conclusion was "long-context agents should move to Fable 5.1." Then I remembered that tokens have to get into the cache before they can be read from it.

Anthropic bills a 5-minute cache write at 1.25x

Anthropic bills a 5-minute cache write at 1.25x the input price for Opus 5. My price table has no separate write price for Fable 5.1. If it follows the same 1.25x rule, writing a 200K prefix costs $1.25 more on Fable 5.1 than on Opus 5. At 200K cached, Fable 5.1 saves $0.015 per request. That's more than 80 requests on the same prefix before the write premium is paid back.

News

Fable 5.1 costs twice as much as Opus 5, until your cache gets big enough

Claude Fable 5.1 shipped on September 1 with the same $10 input / $50 output per million tokens as Fable 5.

@spots #dev
Source: Dev.to
See more like this