Claude Fable 5.1 and Mythos 5.1: faster, half the token usage

Redaktion · · 5 Min. Lesezeit

On September 1, 2026, Anthropic released Claude Fable 5.1, alongside the access-restricted sibling model Claude Mythos 5.1. The API identifier claude-fable-5-1 went live the same day. Per the vendor, the model matches the coding strength of the original Fable but runs faster, uses roughly half the tokens, and communicates more clearly. The base price stays the same — the real cost lever sits in cache reads, which get 75% cheaper.

Where things stood

Fable is Anthropic’s top line built for coding and long agentic runs. With Fable 5, pricing was $10 per 1M input tokens and $50 per 1M output tokens — competitive for a frontier model, but expensive once agentic workflows push large contexts through the model again and again. That is where prompt caching comes in: recurring context blocks are cached so they are billed more cheaply than fresh input on re-use. The cache-read price was therefore a real cost driver for anything reusing the same context multiple times.

In parallel, Fable’s recent headlines were less about the technology and more about access policy — export restrictions around Fable 5 and Mythos 5.

What now applies

1. Half the token usage at the same coding strength. Anthropic states that Fable 5.1 holds the coding performance of the original Fable while needing roughly half the tokens. That is a vendor claim, not an independent benchmark — but if it holds in practice, the lower usage alone cuts real cost per task, regardless of the per-token price.

2. Base price holds, cache reads drop 75%. The price per input and output token is unchanged from Fable 5 ($10 / $50 per 1M). What falls is cache reads: down to $0.25 per 1M tokens, a 75% reduction. Key for costing: this saving applies only to cached context, not to regular input.

3. “Fable for everyone” — less restrictive. Anthropic frames the release as broader availability. TechCrunch calls the version less restrictive than its predecessors. The access-restricted Mythos 5.1 stays separate — it is the more tightly controlled sibling.

Reading

The real cost effect comes from two levers that must be kept apart. The per-token base price stayed the same — anyone using Fable 5.1 just like Fable 5, without caching and without the lower usage, pays nominally the same per token. The savings only materialize through (a) the roughly halved token usage and (b) the 75% cheaper cache reads. Both hit hardest in agentic workflows that reuse the same large context across many steps.

Concretely for costing: the “75% cheaper” figure refers exclusively to cache reads — not to the total price of a call. Booking it in a cost model as a blanket price cut makes you look richer than you are. Realistically, total savings depend heavily on how high the cache-read share is and how large the reused context is in a given workflow.

For agency clients, the direction is still clearly positive: coding-heavy and agentic applications get cheaper to run with Fable 5.1 — provided the setup uses prompt caching consistently. Without caching, the benefit is limited to the lower usage.

What you can do now

If you run Fable in production: check whether your setup actively uses prompt caching. The 75% advantage on cache reads evaporates if recurring contexts are not cached. Measure your workloads’ cache-read share before promising any savings.

If you build cost models: split three line items — input, output, and cache reads. Apply the 75% cut only to the cache-read line, not the total price. Model the halved token usage separately as its own effect.

If you compare models: treat the vendor’s “same coding strength, half the usage” as a hypothesis, not a fact. Run your own comparison against Fable 5 on your real tasks before migrating.

See everything in one place:ClaudeLLM Pricing