Anthropic Claude Prompt Caching: Cost Model and When It Pays Off
August 26, 2026
Anthropic Claude prompt caching cuts read costs 90% but adds a 25% write premium — master the break-even math and decision framework for the CCA-F exam.
Anthropic Claude prompt caching stores reusable content — system prompts, documents, tool definitions — so repeated requests skip re-processing those tokens. Cache writes cost 25% more per token, but cache reads cost 90% less. Whether caching saves money depends on how often the same prefix is reused within the 5-minute default TTL window.
If you have already read Plinth Prep's introduction to cache_control syntax, you know the mechanics. This post covers what that post deliberately leaves out: the economics. Knowing how to place a cache breakpoint is not the same as knowing whether placing one will reduce your bill — and that distinction is exactly what architectural reasoning questions test.
How Much Does Prompt Caching Actually Cost and Save?
Prompt caching changes the cost of input tokens in two directions: a premium to write content into the cache, and a steep discount to read from it. The multipliers below apply to the base input price of whichever model you are using.
| Operation | Cost Multiplier |
|---|---|
| Standard input (no cache) | 1× |
| Cache write — 5-minute TTL | 1.25× |
| Cache write — 1-hour TTL | 2× |
| Cache read | 0.1× |
A Claude Sonnet 4.6 cache write at the 5-minute TTL costs $3.75 per million tokens instead of the standard $3.00. A cache read costs $0.30 per million tokens. The write premium is real and unavoidable on the first request; the read discount is why caching exists.
Worked Example: 10-Request Session on Claude Sonnet 4.6
Assume a shared system prompt of 50,000 tokens — roughly a detailed instruction set plus a loaded document — and 10 requests within a single session, all reusing that prefix within the 5-minute window.
Without caching: 10 × 50,000 tokens × $3.00 per million = $1.50
With caching at the 5-minute TTL:
- First request writes to cache: 50,000 × $3.75/M = $0.19
- Nine cache reads: 9 × 50,000 × $0.30/M = $0.14
- Total: $0.33
That is a roughly 78% reduction on session input cost. Across thousands of sessions per day the savings compound rapidly. However, the same math also reveals the failure mode: if only one request hits that prefix, you pay $0.19 versus $0.15 uncached — caching made it more expensive.
Break-Even by TTL
With the 5-minute TTL, two requests break even: one write at 1.25× plus one read at 0.1× totals 1.35×, which is less than two uncached requests at 2×. With the 1-hour TTL, the 2× write cost means you need at least three requests before you come out ahead (2× write + 2 reads at 0.1× = 2.2×, versus three uncached at 3×).
The practical rule: if an endpoint receives fewer than two requests against the same prefix per TTL window, caching costs more than it saves at the 5-minute tier. Fewer than three, and the 1-hour TTL is not worth its higher write cost.
Which Claude Models Support Prompt Caching?
All current Claude models support prompt caching: Opus 4.8, Opus 4.7, Opus 4.6, Sonnet 4.6, and Haiku 4.5. Each model enforces a minimum cacheable prefix — prefixes below this threshold are silently skipped. No error is returned; the cache is simply never written, and cache_creation_input_tokens in the usage object will be zero.
- Opus 4.8, Opus 4.7, Opus 4.6, and Haiku 4.5: 4,096 tokens minimum
- Sonnet 4.6: 2,048 tokens minimum
This asymmetry matters architecturally. A 3,000-token system prompt will cache on Sonnet 4.6 but silently fail to cache on Opus 4.8 — despite both models accepting the same cache_control marker. When debugging a cache miss, verify the prefix length against the model's minimum before investigating more complex invalidation causes.
When Does Prompt Caching Pay Off — and When Does It Not?
Caching earns its keep in three situations:
- Large, stable shared prefixes. Document Q&A systems, coding assistants loaded with repository context, and support bots with lengthy policy documents all have the right profile: the prefix is big enough to clear the minimum threshold and reused frequently enough to generate many reads per write.
- Multi-turn conversations. Placing a breakpoint at the end of the most-recently-appended turn lets each subsequent response reuse the growing conversation prefix, with cache hits accruing incrementally across the session.
- Shared tool definitions. Tool schemas render before system content. A large tool library caches with a single breakpoint at the end of the tool list, and every request in the session benefits.
Caching costs more — or does nothing — in three situations:
- Dynamic content injected into the stable prefix. A timestamp, UUID, or per-user identifier interpolated into the system prompt changes the prefix on every request. Each call pays the write premium; none trigger a read. This is the single most common cause of unexpected cost increases when caching is enabled.
- Low-volume or singleton endpoints. Any route that processes one request per session and never revisits the same prefix pays the write premium with no offsetting reads.
- Prefixes below the minimum threshold. A 1,500-token system prompt on any Opus model never caches. The
cache_controlmarker is accepted without error but has no effect.
How Do Cache TTLs Shape Your Architecture?
The 5-minute default TTL suits continuous, high-frequency workloads. If real requests arrive faster than the cache would expire, those requests keep it warm automatically — no dedicated warm-up step is needed.
The 1-hour TTL suits bursty traffic patterns: a cluster of requests, a 20-minute gap, another cluster. Without the longer TTL, every burst would pay a fresh write premium. With it, you pay the 2× write cost once and amortize it across the burst.
The choice collapses to a single calculation: if your inter-request gaps routinely exceed 5 minutes and your bursts contain at least three requests, the 1-hour TTL reduces total write charges. If gaps are short or bursts are small, the default is cheaper.
One operational consequence worth knowing: switching models mid-session invalidates the cache. Caches are keyed to the model ID. Swapping from Sonnet 4.6 to Opus 4.8 discards all entries and forces a fresh write at the new model's rate. For the same reason, reordering tool definitions or changing any byte of the stable prefix before a breakpoint resets the cache for everything downstream.
What CCA-F Candidates Need to Know About Caching Trade-offs
Based on Anthropic's public documentation, candidates should expect reasoning questions that describe an architecture and ask whether caching will reduce cost, increase cost, or have no effect. Three concepts anchor most of those questions:
- Cost direction is not obvious at request one. On the very first request, caching always costs more. The question is whether the downstream reads justify the write premium. An architect who treats caching as universally beneficial is wrong; so is one who treats it as universally wasteful.
- Silent invalidation is an architecture problem, not a configuration problem. Adding
cache_controlto a prefix that contains adatetime.now()call does not cause an error — it causes the cache to be rewritten on every request while appearing to work. Diagnosing this requires readingcache_read_input_tokens, not just confirming the marker is present. - TTL selection follows traffic shape, not instinct. Choosing the 1-hour TTL because "longer is better" is an error. The 1-hour TTL costs more per write and only pays off under specific traffic conditions. An exam question that asks you to recommend a TTL for a low-volume endpoint has a defensible answer: the default 5-minute TTL, or no caching at all.
Quick Reference: Prompt Caching Decision Checklist
- ☐ Is the shared prefix at least 4,096 tokens (Opus 4.x or Haiku 4.5) or 2,048 tokens (Sonnet 4.6)?
- ☐ Will at least two requests hit the same prefix within the TTL window (five-minute tier) — or at least three (one-hour tier)?
- ☐ Is the prefix free of dynamic content — no timestamps, UUIDs, or per-request identifiers before the breakpoint?
- ☐ Are tool definitions serialized deterministically — sorted, not randomly ordered — so the rendered bytes are stable?
- ☐ Will the same model be used for the duration of the session?
- ☐ If traffic is bursty with gaps longer than five minutes, is the 1-hour TTL worth its 2× write cost given expected burst size?
Every unchecked box is a potential reason cache_read_input_tokens returns zero. Address the underlying condition before adding breakpoints — otherwise you pay the write premium without collecting the read discount that justifies it.
Frequently asked questions
- What does prompt caching cost in Claude?
- Cache writes cost 25% more than standard input tokens (1.25× the base rate) at the default 5-minute TTL, or 2× for the 1-hour TTL. Cache reads cost 90% less (0.1× the base rate). On Claude Sonnet 4.6, a cached read costs $0.30 per million tokens versus $3.00 uncached.
- How many requests does it take to break even with Claude prompt caching?
- With the default 5-minute TTL, two requests break even: one cache write at 1.25× and one read at 0.1× totals 1.35× — less than two uncached requests at 2×. The 1-hour TTL requires at least three requests to recoup its higher 2× write cost.
- Which Claude models support prompt caching?
- All current Claude models support prompt caching: Opus 4.8, Opus 4.7, Opus 4.6, Sonnet 4.6, and Haiku 4.5. The minimum cacheable prefix is 4,096 tokens for Opus and Haiku 4.5 models, and 2,048 tokens for Sonnet 4.6. Prefixes below the minimum are silently skipped with no error returned.
- When does prompt caching not save money?
- Prompt caching costs more when the cached prefix changes between requests — a date, UUID, or per-session ID in the system prompt invalidates the cache on every call. It also costs more when request volume is too low to offset the write premium, or when the prefix is shorter than the model's minimum cacheable threshold.
- What is the TTL for Claude prompt caching and how does it affect cost?
- Claude prompt caching offers two TTL options: 5 minutes (the default, billed at 1.25× base write rate) and 1 hour (billed at 2× base write rate). The 1-hour option suits bursty traffic with gaps longer than five minutes; the 5-minute default is cheaper for continuous high-frequency workloads with short inter-request gaps.