Every production AI request re-sends text the provider has already read: the same system prompt, the same tool definitions, the same document, the same forty turns of conversation. Prompt caching lets the provider keep the processed state of that repeated beginning — the prefix — and bill you a fraction of the input price for reusing it. On all four major APIs a cache read now costs a tenth of fresh input or less.
What differs is everything around the read. Anthropic and OpenAI's newest models charge 1.25× input to write a prefix into the cache; Google and DeepSeek charge nothing extra. Anthropic only caches when you ask; OpenAI, Google and DeepSeek cache by default. The minimum prefix ranges from none to 4,096 tokens, and the lifetime from five minutes to "a few hours to a few days". Those details decide whether caching cuts a bill by 80% or quietly raises it.
We read each provider's pricing page and caching documentation on September 29, 2026 and priced three real workload shapes on five models. The short answer: caching cut the modeled bills by 67–89% where most of the input repeats, by 15–17% where only the instructions repeat — and made one bill 20% more expensive than no caching at all when a timestamp broke the prefix.
The four rulebooksSame idea, different fine print
Anthropic (Claude)
- How it turns on
- You mark it: one top-level cache_control, or up to 4 breakpoints
- Write
- 1.25× (5 min) · 2× (1 hour)
- Read
- 0.1× · 0.05× Opus 5.5 · 0.025× Fable 5.1
- Minimum prefix
- 512 tokens on the newest models; 4,096 on Haiku 4.5
- Lifetime
- 5 min or 1 hour, refreshed by every read
OpenAI (GPT-5.6, GPT-6)
- How it turns on
- On by default; implicit or explicit breakpoints
- Write
- 1.25×
- Read
- 0.1×
- Minimum prefix
- 1,024 tokens
- Lifetime
- 30 minutes, refreshed by every reuse
OpenAI (GPT-5.5 and older)
- How it turns on
- On by default, implicit only
- Write
- no fee
- Read
- model-dependent (0.1× on GPT-5.5)
- Minimum prefix
- varies by request settings
- Lifetime
- 5–10 min idle, or up to 24 hours
Google (Gemini 2.5 and newer)
- How it turns on
- Implicit by default; explicit cache objects optional
- Write
- no fee (explicit: storage per hour)
- Read
- 0.1×
- Minimum prefix
- 4,096 tokens on Gemini 3.x; 2,048 on 2.5
- Lifetime
- implicit: not stated; explicit: the TTL you set
DeepSeek (V4.1 Flash, V4 Pro)
- How it turns on
- Always on, disk cache
- Write
- no fee
- Read
- 0.02× Flash · 0.033× Pro
- Minimum prefix
- none stated
- Lifetime
- hours to days, best effort
| Provider | How it turns on | Write | Read | Minimum prefix | Lifetime |
|---|---|---|---|---|---|
| Anthropic (Claude) | You mark it: one top-level cache_control, or up to 4 breakpoints | 1.25× (5 min) · 2× (1 hour) | 0.1× · 0.05× Opus 5.5 · 0.025× Fable 5.1 | 512 tokens on the newest models; 4,096 on Haiku 4.5 | 5 min or 1 hour, refreshed by every read |
| OpenAI (GPT-5.6, GPT-6) | On by default; implicit or explicit breakpoints | 1.25× | 0.1× | 1,024 tokens | 30 minutes, refreshed by every reuse |
| OpenAI (GPT-5.5 and older) | On by default, implicit only | no fee | model-dependent (0.1× on GPT-5.5) | varies by request settings | 5–10 min idle, or up to 24 hours |
| Google (Gemini 2.5 and newer) | Implicit by default; explicit cache objects optional | no fee (explicit: storage per hour) | 0.1× | 4,096 tokens on Gemini 3.x; 2,048 on 2.5 | implicit: not stated; explicit: the TTL you set |
| DeepSeek (V4.1 Flash, V4 Pro) | Always on, disk cache | no fee | 0.02× Flash · 0.033× Pro | none stated | hours to days, best effort |
Source: CalculatorAI · calculatorai.app · platform.claude.com · developers.openai.com · ai.google.dev · api-docs.deepseek.com
Four things in that table matter more than the rest.
The write fee is new at OpenAI. Since GPT-5.6, OpenAI bills a cache write at 1.25× input, like Anthropic; the older GPT-5.5, GPT-5.4 and GPT-5 families still cache for free. The rate cards themselves are decoded line by line in our OpenAI API pricing guide and Claude API pricing guide. Claude Sonnet 5.5 and GPT-6 Sol now have identical cache prices — $2 input, $2.50 write, $0.20 read, $10 output per million — and differ only in the minimum (512 vs 1,024 tokens) and the lifetime (5 minutes or an hour vs 30 minutes).
Anthropic is the only one that does nothing unless asked. A Claude request without cache_control is never cached. The upside is control: nothing is written that you did not mark. On OpenAI's GPT-5.6+ caching is "enabled by default", and in implicit mode the breakpoint goes at the end of the latest user message — which, as the RAG example below shows, can mean paying a write on text nobody will ever read again.
Opus 5.5 reads its cache at Sonnet's price. Anthropic cut the read multiplier on Claude Opus 5.5 to 0.05×, so a cached Opus 5.5 prefix costs $0.20 per million — the same as Sonnet 5.5, although fresh Opus input is twice the price. Claude Fable 5.1 reads at 0.025×. The more expensive the model, the more a cache hit is worth.
Google's floor is high. Implicit caching on Gemini 3.x only applies to prompts of at least 4,096 tokens. A 2,000-token system prompt gets no discount on Gemini at any volume. Gemini also offers explicit cache objects, which carry a storage charge per hour on top of the discounted read — the trade-off our Gemini API pricing guide prices for a document workload.
The break-evenWhen a write pays for itself
Caching changes the price of the prefix only. For a prefix sent N times, the arithmetic is the same on every provider — only the two multipliers change:
prefix × input price × Nprefix × input price × (write + read × (N − 1))1.25 + 0.1 × (N − 1) — cheaper from the 2nd use2 + 0.1 × (N − 1) — cheaper from the 3rd use1 + read × (N − 1) — never more expensiveEvery setup saves 71–88% on a prefix reused ten times; the differences only matter at low reuse.
Show these figures as a table
| Value (% of the uncached prefix cost) | |
|---|---|
| Anthropic 1-hour (Sonnet 5.5) — 2× write | 29 |
| Anthropic 5-minute / OpenAI GPT-6 — 1.25× write, 0.1× read | 21.5 |
| Gemini implicit — no write fee | 19 |
| Anthropic 5-minute (Opus 5.5) — 0.05× read | 17 |
| DeepSeek V4.1 Flash — 0.02× read | 11.8 |
Source: CalculatorAI · calculatorai.app · Provider pricing pages, September 29, 2026 · drafts/prompt-caching-numbers.mjs
The lesson of the break-even is not "caching is always right". It is that a write that is never read is a 25% surcharge on those tokens — on Anthropic and on OpenAI's newest models. Everything that follows is about making sure the tokens you write are the ones you read.
Workload 1A support bot with a fixed prompt
The simplest case: a 6,000-token system prompt (instructions, policies, product FAQ), a 300-token customer message and a 250-token answer, 50,000 replies a month. The only question is the hit rate — the share of requests that find the prefix still cached. A miss re-writes it.
Claude Opus 5.5
- No caching
- $1,510
- 50% hits
- $1,090
- 80% hits
- $658
- 95% hits
- $442
Claude Sonnet 5.5 / GPT-6 Sol
- No caching
- $755
- 50% hits
- $560
- 80% hits
- $353
- 95% hits
- $250
Gemini 3.8 Flash
- No caching
- $283
- 50% hits
- $182
- 80% hits
- $121
- 95% hits
- $91
DeepSeek V4.1 Flash
- No caching
- $55
- 50% hits
- $33
- 80% hits
- $19
- 95% hits
- $13
| Model | No caching | 50% hits | 80% hits | 95% hits |
|---|---|---|---|---|
| Claude Opus 5.5 | $1,510 | $1,090 | $658 | $442 |
| Claude Sonnet 5.5 / GPT-6 Sol | $755 | $560 | $353 | $250 |
| Gemini 3.8 Flash | $283 | $182 | $121 | $91 |
| DeepSeek V4.1 Flash | $55 | $33 | $19 | $13 |
Source: CalculatorAI · calculatorai.app · drafts/prompt-caching-numbers.mjs
At a 95% hit rate caching takes 67–77% off the whole bill, not just the prefix, because the prefix is 95% of every request's input. And the hit rate is mostly a matter of traffic: 50,000 replies a month is more than one a minute, so during working hours each request arrives well inside a 5-minute (Claude) or 30-minute (OpenAI) window and refreshes the entry for the next one. Misses cluster at the start of the day and after quiet spells.
Workload 2A coding agent whose context grows
Agents are where caching stops being an optimisation and becomes the price. Each step of an agent loop re-sends the entire conversation so far — system prompt, tool definitions, files read, every tool result — and appends a little more. We modeled a 30-step session starting from 30,000 tokens (instructions, tools, a repository map), adding 2,500 tokens of tool results and 700 tokens of model output per step. By the last step the context is 122,800 tokens, and the session has sent 2.29 million input tokens in total — almost all of them repeats.
With caching, the breakpoint moves to the end of each step: the next step reads everything already cached and writes only the new increment.
30,000-token start, +3,200 tokens per step, 700 tokens written per step
Without caching the session pays full input price on 2.29 million tokens. With the breakpoint moved to the end of each step, each step writes only its new ~3,200 tokens and reads the rest. At 200 sessions a month the difference on Sonnet 5.5 or GPT-6 Sol is $959 against $190.
Two things follow. First, the model choice and the caching choice are the same size of decision: an uncached Sonnet session costs more than a cached Opus session. Second, anything that changes the start of the context mid-session — reordering tools, editing the system prompt, switching model — throws away the whole cache and re-writes 100,000+ tokens at 1.25×. Tool definitions are usually the bulk of that start; our MCP token cost analysis shows servers whose definitions are 99% of the input.
Workload 3RAG, where only the instructions repeat
Retrieval-augmented generation is the case where caching disappoints, and where it can make things worse. A typical request: 2,000 tokens of stable instructions, 8,000 tokens of retrieved passages that are different for every question, a 100-token question and a 400-token answer, 20,000 questions a month. Only the first 2,000 tokens ever repeat.
$484 → $412 a month (−15%)
Sonnet 5.5 or GPT-6 Sol. The 2,000 instruction tokens are read from cache; the 8,100 unique tokens are sent at normal input price. Caching saves little because little repeats — but every dollar it touches is saved.
$484 → $493 a month (+2%)
Same model and traffic, but the cache is written through the unique passages and question — what automatic caching on Claude or implicit mode on GPT-5.6+ does when the prompt ends in content that never repeats. The instructions are still read at 0.1×, but 8,100 tokens per request are written at 1.25× and never read.
Anthropic's documentation names this pattern explicitly: when a prompt ends in per-request content, put an explicit breakpoint at the end of the shared part. OpenAI's guide gives the same advice for GPT-5.6+ and adds an explicit-only mode, in which nothing is written unless you mark it.
The other two providers fail differently. Gemini 3.8 Flash gives this workload no discount at all, because 2,000 tokens is below its 4,096-token implicit minimum — the bill is $181.50 however many questions share the instructions. Growing the instructions to 4,096 tokens (with worked examples the model can actually use) makes each request longer but cheaper: $157.64 a month. DeepSeek detects a prefix shared across requests and persists it after it has seen it, so the instructions hit from roughly the third request on: $35.10 becomes $29.22 at off-peak rates, a 17% cut.
What silently breaks a cacheFive common causes
A cache only matches a byte-identical prefix. Nothing warns you when it stops matching: requests keep succeeding, and the bill goes up.
A timestamp or ID in the system prompt
"Today is 2026-09-29 14:03" changes every minute, so every request writes a new prefix and none reads one. On the support bot above with Claude's automatic caching, that turns a $250 month into $905 — 20% more than not caching at all ($755). Put the date after the last breakpoint, or at day resolution.
Tools or JSON in a different order
Tool definitions render before the system prompt on Claude, so a tool list built from an unordered set, or a JSON schema with keys in a different order, changes the start of every request. Sort them.
Switching model or effort mid-conversation
Caches are per model. Routing a follow-up to a cheaper model starts a fresh cache there; on Claude, changing top-level effort also invalidates the cached messages.
Editing a message instead of appending one
OpenAI's guide notes that extending the last message (Content A becoming A + B) moves the old breakpoint inside a message, so the saved prefix is not reused. Append a new message instead.
A prefix under the minimum
Below 512 tokens on current Claude models, 1,024 on GPT-5.6+, 4,096 on Gemini 3.x, nothing is cached and nothing errors. On Claude the tell is cache_creation_input_tokens: 0; OpenAI's own guide shows a case where padding a short prefix to 1,024 tokens is cheaper after enough requests.
StackingCaching plus Batch
Anthropic's pricing page states that its cache multipliers stack with the Batch API's 50% discount. Run the support-bot workload through Batch with the same 95% hit rate and the Sonnet 5.5 bill falls from $755 uncached in real time to $125 — a sixth. The catch is on the hit rate: requests in a batch are processed independently, so how many read the cache depends on scheduling, and the figure is a best case. The same "cache first, batch second" order is priced on a single-document job in our Claude API pricing guide. OpenAI and Google offer their own Batch tiers; DeepSeek's equivalent lever is the clock — its off-peak rates are half the peak rates.
How to see your own hit rateThe usage fields
Every provider reports cached tokens in the response, so the hit rate is measured, not guessed:
Anthropic
usage.cache_read_input_tokens (read), usage.cache_creation_input_tokens (written), usage.input_tokens (the uncached tail). Reads at zero on repeated requests mean something in the prefix is changing.
OpenAI
usage.input_tokens_details.cached_tokens and, on GPT-5.6+, cache_write_tokens. On older models cached_tokens is rounded down to a multiple of 128.
Google Gemini
The cached token count in the response's usage metadata. Explicit caches also appear as storage on the bill.
DeepSeek
usage.prompt_cache_hit_tokens and prompt_cache_miss_tokens. The cache is best-effort, so measure rather than assume a rate.
Divide cached tokens by all input tokens over a day of real traffic and you have the number every table in this article turns on. The AI Token Cost Calculator takes that cached share as an input and prices the request on any of these models; for a loop whose context grows, the AI Agent Cost Calculator models the caching step by step. If DeepSeek's near-free cache reads tempt you, the whole-task comparison is in DeepSeek vs OpenAI API cost.
Where these numbers come from
Prices and caching rules were read on September 29, 2026 from Anthropic's pricing page and prompt-caching documentation (platform.claude.com), OpenAI's pricing page and prompt-caching guide (developers.openai.com), Google's Gemini API pricing and context-caching pages (ai.google.dev), and DeepSeek's pricing and context-caching pages (api-docs.deepseek.com). Models and per-million rates used: Claude Sonnet 5.5 $2 in / $2.50 5-minute write / $0.20 read / $10 out; Claude Opus 5.5 $4 / $5 / $0.20 / $20; GPT-6 Sol $2 / $2.50 / $0.20 / $10; Gemini 3.8 Flash $0.75 / no write / $0.075 / $3.75 (its 2026 rate; Google lists $1.50 / $0.15 / $7.50 from January 1, 2027); DeepSeek V4.1 Flash off-peak $0.15 miss / $0.003 hit / $0.60 out (peak rates are double). All figures are arithmetic in drafts/prompt-caching-numbers.mjs on stated token counts. Assumptions and their bias: the same token counts are used on every provider, although tokenizers differ (Anthropic says Claude 4.7 and later produce about 30% more tokens for the same text), so compare the percentage savings across providers more than the dollar levels; the agent model assumes every step lands inside the cache lifetime and pays exactly one write per step, which favours caching; the support-bot hit rates are inputs, not measurements; Gemini's explicit caching and its storage charge are not modeled in the workloads; DeepSeek's cache is best-effort and its RAG figure assumes the shared prefix hits from the start, which slightly overstates the saving. Prices change; the method does not.
Frequently asked questions
What is prompt caching? It is a provider-side store of the processed beginning of a prompt. When a later request starts with the byte-identical text, the provider reuses the stored state instead of reprocessing it, and bills those tokens at a reduced "cached input" rate — a tenth of the normal input price or less on Anthropic, OpenAI, Google and DeepSeek.
How much does prompt caching save? On the prefix itself, 71–88% when it is reused ten times. On a whole bill it depends on how much of each request repeats: 67–77% for a support bot with a 6,000-token fixed prompt at a 95% hit rate, 80–89% for a 30-step coding agent, and only 15–17% for a RAG system where just the instructions repeat.
Does prompt caching cost extra? On Anthropic and on OpenAI's GPT-5.6 and GPT-6 models, writing a prefix into the cache costs 1.25× the input price (2× for Claude's 1-hour cache), so a cached prefix that is never reused costs more than an uncached one. Google's implicit caching and DeepSeek's cache charge nothing to write; Google's explicit caches charge storage per hour.
Is prompt caching automatic? On OpenAI, Google Gemini (2.5 and newer) and DeepSeek it is on by default. On Claude it is opt-in: add a top-level cache_control field for automatic placement, or mark up to four breakpoints yourself.
How long does a cached prompt last? Five minutes or one hour on Claude, refreshed by every read; 30 minutes on OpenAI's GPT-5.6 and later, refreshed by every reuse; from 5–10 minutes of inactivity up to 24 hours on older OpenAI models; the time-to-live you set for Gemini's explicit caches; and "a few hours to a few days", best effort, on DeepSeek.
Why is my cache hit rate zero? Usually because something at the start of the prompt changes on every request — a timestamp, a request ID, tools in a different order — or because the prefix is shorter than the model's minimum (512 tokens on current Claude models, 1,024 on GPT-5.6+, 4,096 on Gemini 3.x). Neither produces an error; the cached-token field in the response simply stays at zero.






