Prompt Caching for LLM APIs — Mechanics, Cost and TTL Choice

Martin Rau ·

A system prompt with tool definitions, repository context, or a long knowledge base repeats almost word-for-word on every call in a conversation. Without a countermeasure, that block gets paid for again on every single request — in tokens and in latency. Prompt caching fixes exactly that: a provider stores the internal model state after a stable prompt section and reuses it on the next request instead of recomputing it. The glossary entry on prompt caching covers the short definition — this article goes into the mechanics: how caching actually kicks in, what it costs, where models differ, and when it genuinely pays off in agent and API workloads.

How prompt caching works technically

Prompt caching is a pure prefix match. The provider computes the internal key-value states for a prompt section once and stores them under a hash of the exact bytes. On the next request, it compares the beginning of the new prompt byte-for-byte against the cached prefix — if it matches, the stored states are reused, and only the part after it needs fresh processing.

Two details decide whether this actually kicks in:

  • Order matters. Requests are assembled from tool definitions, system prompt, and message history — rendered in exactly that order. Whatever sits earlier in the prompt must be more stable than what follows. A single changed byte makes everything after that point uncacheable, no matter how many cache markers are set.
  • Cache markers (cache_control) set the breakpoints. With Anthropic, {"type": "ephemeral"} explicitly marks how far a block should be cached — up to four such breakpoints per request. OpenAI and Google don’t require explicit markers for this (see below), but follow the same prefix principle.

Classic setup for an agent loop: one breakpoint at the end of the static system prompt (role, tool descriptions, fixed instructions), a second at the end of the growing conversation history. Volatile content — timestamps, random IDs, the current user question — belongs after that, never before it.

What prompt caching costs

Cache hits are noticeably cheaper than the regular input price, but creating a cache entry itself carries a surcharge — the math only pays off after a certain amount of reuse.

At Anthropic, the following currently applies (as of September 2026):

  • Writing a cache (the first request that stores the block anew): 1.25× the regular input price at the 5-minute TTL, 2× at the 1-hour TTL.
  • Reading a cache (a hit on a follow-up request): roughly 0.1× the regular input price.

That gives a simple break-even: at the 5-minute TTL, the write surcharge already pays off from the second request sharing the same prefix (1.25× + 0.1× ≈ 1.35× versus 2× uncached for two requests). At the 1-hour TTL, you need at least three requests, because the doubled write price requires more hits to balance out. So caching doesn’t pay off for a single, isolated request — only once the same prompt prefix gets reused multiple times, as in a multi-turn dialogue or repeated, orchestrated calls that share a system prompt.

TTL variants: 5 minutes vs. 1 hour

Anthropic offers two lifetimes for a cache entry, chosen by the time gap between requests that share the same prefix:

| Gap between requests | Recommended TTL | |---|---| | Under 5 minutes (continuous traffic, fast agent turns) | 5 minutes (default) — every request refreshes the entry on its own | | 5 to 60 minutes (a user replies after a pause, a longer background task) | 1 hour — this is where the doubled write price actually pays off | | Over 1 hour | Neither TTL helps directly — pre-warm the cache deliberately or accept the miss |

Important detail: a cache read extends the lifetime for free, but the clock runs from the start of the request — not after the response finishes. A generation that takes four minutes leaves only about one minute for the next request under a 5-minute TTL.

Model differences: minimum length for a cache hit

A prompt section has to reach a minimum length before it gets cached at all — shorter prefixes are silently ignored, with no error, recognizable only by the cache-write value staying at zero. At Anthropic, this minimum is model-dependent and not consistently tiered:

| Model generation (examples) | Minimum | |---|---:| | Claude Opus 5 (newest generation) | 512 tokens | | Claude Sonnet 5 | 1,024 tokens | | Older intermediate generations | 2,048 tokens | | Some smaller/older models | 4,096 tokens |

A short system prompt of around 700 tokens caches without issue on Claude Opus 5, but stays ineffective on a model with a 4,096-token minimum — without the call itself failing. Anyone switching between models should check this value for the provider in question rather than carrying it over from an earlier model generation.

When prompt caching genuinely pays off

Four situations where caching measurably pays for itself in everyday use:

  1. Multi-turn agent dialogues. A coding agent with 30,000 tokens of repository context only pays the cache-read price for that context after the first call — the next twenty conversation turns cost full price only for the tokens that are genuinely new.
  2. Orchestrated calls with a fixed system prompt. Several worker agents using the same system prompt and the same tool definitions share the cache entry — as long as the requests come from the same workspace and the prefix stays byte-identical.
  3. Repeated calls against the same knowledge base. A long reference document serving as system context for many individual questions (support bot, document Q&A) benefits strongly, because the document itself gets cached and only the individual question is processed fresh.
  4. Batch-like tasks with a shared prefix. Classification or extraction across many records with identical instructions in the system prompt — here caching pays off after only a few repetitions, provided the TTL matches the call frequency.

Conversely, caching brings nothing for genuinely one-off requests without repetition, for prompts that differ completely every time, or when a dynamic component (timestamp, session ID) accidentally ends up before the breakpoint.

Distinction: how other providers cache

Not every provider implements prompt caching the same way — the differences are mainly about automation and cost model:

  • OpenAI caches automatically, without needing explicit cache_control markers — starting at a prompt length of 1,024 tokens, in further increments of 128 tokens. This also applies to models like GPT-6 Astra. There is no separate write surcharge like at Anthropic; cache hits are simply billed at a discounted rate.
  • Google Gemini runs two mechanisms in parallel: implicit caching happens automatically in the background (for example on models like Gemini 3 Pro) and lowers the price on a hit, with no guarantee that a hit occurs at all. Explicit caching, by contrast, actively creates a cache entry with a cost guarantee — but adds a time-based storage fee for the entry’s lifetime, regardless of whether it gets read.

For practical work, this means: with Anthropic you control breakpoints and TTL yourself and pay a write surcharge that only pays off after several hits. With OpenAI, caching happens automatically in the background with no extra cost for creating it. With Google, you have to choose between automatic, uncertain caching and explicit, but time-billed, caching.

FAQ

Do I have to activate prompt caching manually? At Anthropic, yes — via cache_control markers at the desired breakpoints, otherwise nothing gets cached. At OpenAI it runs automatically once the minimum length is reached, with no action needed. Google offers both: automatic (implicit) and manual (explicit, with a cost guarantee).

How do I tell whether a cache hit actually happened? From the response’s usage fields: one for newly written cache tokens and one for tokens read from cache. If the read value stays at zero across several requests sharing the same prefix, a silent invalidation issue is at work — usually a dynamic element sitting before the breakpoint.

Does the 1-hour TTL pay off compared to the 5-minute variant? Only for gaps between 5 and 60 minutes between requests sharing the same prefix. Under continuous traffic below 5 minutes, the cheaper 5-minute cache renews itself anyway — the more expensive hourly variant then only adds the doubled write price with no extra benefit.

Why doesn’t my short system prompt cache? Every model has a minimum prefix length before a cache entry gets created at all — it ranges from a few hundred to several thousand tokens depending on the model. Below that threshold, the call stays technically error-free but simply doesn’t cache.

What’s the difference between this lexicon article and the glossary entry? The glossary entry gives the compact definition in a few sentences. This article goes into the mechanics: the cost model with concrete multipliers, TTL choice, model-dependent minimum lengths, and the distinction from other providers.

Topic overview