Every provider publishes a price per million tokens, and every comparison article lines those prices up in a table. That table answers a question nobody has. You do not buy a million tokens; you buy a summary, a support reply, a translation, a coding-agent turn, a conversation. Each of those has its own shape — how much goes in, how much of it the provider has already seen, how much comes out — and the shape decides the bill far more than the headline rate does.
So this comparison prices five real tasks on nine models: OpenAI, Anthropic and Google, each at three tiers (their flagship, their mid-tier workhorse, their cheap model). Every rate comes from the same verified price table the AI Cost Calculator uses, read from the providers' own pricing pages on September 9, 2026. The arithmetic is in the open, and you can change any assumption in the calculator and get your own numbers.
Short version: at the same tier, Gemini is the cheapest of the three on every task, usually by about 3×; OpenAI's cheap model is the cheapest thing on the list; and which tier you choose matters ten times more than which provider. A flagship model costs 10–20× its own cheap sibling on identical work. That gap, not the gap between logos, is where an AI bill is won or lost.
The nine modelsSame tiers, three price lists
Each provider sells a flagship for hard reasoning, a mid-tier model that most production work should run on, and a small fast model for simple, high-volume jobs. Lining those up is the only fair way to compare, because a flagship against a competitor's cheap model tells you nothing except that cheap is cheaper.
Flagship
- Model
- OpenAI GPT-5.6 Sol
- Input
- $4.00
- Cached input
- $0.40
- Output
- $20.00
Flagship
- Model
- Anthropic Claude Opus 5
- Input
- $5.00
- Cached input
- $0.50
- Output
- $25.00
Flagship
- Model
- Google Gemini 3.1 Pro
- Input
- $2.00
- Cached input
- $0.20
- Output
- $12.00
Mid-tier
- Model
- OpenAI GPT-5.6 Terra
- Input
- $2.00
- Cached input
- $0.20
- Output
- $12.00
Mid-tier
- Model
- Anthropic Claude Sonnet 5
- Input
- $2.00
- Cached input
- $0.20
- Output
- $10.00
Mid-tier
- Model
- Google Gemini 3.8 Flash
- Input
- $0.75
- Cached input
- $0.075
- Output
- $3.75
Cheap
- Model
- OpenAI GPT-5.6 Luna
- Input
- $0.20
- Cached input
- $0.02
- Output
- $1.20
Cheap
- Model
- Anthropic Claude Haiku 4.5
- Input
- $1.00
- Cached input
- $0.10
- Output
- $5.00
Cheap
- Model
- Google Gemini 3.5 Flash-Lite
- Input
- $0.30
- Cached input
- $0.03
- Output
- $2.50
| Tier | Model | Input | Cached input | Output |
|---|---|---|---|---|
| Flagship | OpenAI GPT-5.6 Sol | $4.00 | $0.40 | $20.00 |
| Flagship | Anthropic Claude Opus 5 | $5.00 | $0.50 | $25.00 |
| Flagship | Google Gemini 3.1 Pro | $2.00 | $0.20 | $12.00 |
| Mid-tier | OpenAI GPT-5.6 Terra | $2.00 | $0.20 | $12.00 |
| Mid-tier | Anthropic Claude Sonnet 5 | $2.00 | $0.20 | $10.00 |
| Mid-tier | Google Gemini 3.8 Flash | $0.75 | $0.075 | $3.75 |
| Cheap | OpenAI GPT-5.6 Luna | $0.20 | $0.02 | $1.20 |
| Cheap | Anthropic Claude Haiku 4.5 | $1.00 | $0.10 | $5.00 |
| Cheap | Google Gemini 3.5 Flash-Lite | $0.30 | $0.03 | $2.50 |
Source: CalculatorAI · calculatorai.app · Provider pricing pages via the CalculatorAI price table (src/config/ai-models.ts)
Two things in that table already decide most of what follows. Output is 5–6× the price of input everywhere, so a task that writes a lot costs more than a task that reads a lot, even when the reading is longer. And cached input is a tenth of base input, so any task with a fixed prefix — a system prompt, a document, a codebase — is mostly a question of how much of it the cache absorbs.
Five tasks, pricedWhat one unit of work costs
The tasks are deliberately ordinary. For each one the token shape is fixed, so the only thing that changes across the columns is the price list.
Summarise a 20-page PDF
- OpenAI
- $0.076
- Anthropic
- $0.123
- $0.040
Support reply, cached 3k system prompt
- OpenAI
- $0.0093
- Anthropic
- $0.015
- $0.0052
Translate 1,000 words
- OpenAI
- $0.038
- Anthropic
- $0.062
- $0.022
Coding-agent turn, 60k context, 90% cached
- OpenAI
- $0.086
- Anthropic
- $0.139
- $0.047
30-turn chat, growing context, cached
- OpenAI
- $0.505
- Anthropic
- $0.821
- $0.277
| Task | OpenAI | Anthropic | |
|---|---|---|---|
| Summarise a 20-page PDF | $0.076 | $0.123 | $0.040 |
| Support reply, cached 3k system prompt | $0.0093 | $0.015 | $0.0052 |
| Translate 1,000 words | $0.038 | $0.062 | $0.022 |
| Coding-agent turn, 60k context, 90% cached | $0.086 | $0.139 | $0.047 |
| 30-turn chat, growing context, cached | $0.505 | $0.821 | $0.277 |
Source: CalculatorAI · calculatorai.app · drafts/chatgpt-vs-claude-vs-gemini-cost-per-task-numbers.mts — September 2026 rates
Summarise a 20-page PDF
- OpenAI
- $0.040
- Anthropic
- $0.049
- $0.014
Support reply, cached 3k system prompt
- OpenAI
- $0.0052
- Anthropic
- $0.0060
- $0.0017
Translate 1,000 words
- OpenAI
- $0.022
- Anthropic
- $0.025
- $0.0071
Coding-agent turn, 60k context, 90% cached
- OpenAI
- $0.047
- Anthropic
- $0.056
- $0.016
30-turn chat, growing context, cached
- OpenAI
- $0.277
- Anthropic
- $0.329
- $0.095
| Task | OpenAI | Anthropic | |
|---|---|---|---|
| Summarise a 20-page PDF | $0.040 | $0.049 | $0.014 |
| Support reply, cached 3k system prompt | $0.0052 | $0.0060 | $0.0017 |
| Translate 1,000 words | $0.022 | $0.025 | $0.0071 |
| Coding-agent turn, 60k context, 90% cached | $0.047 | $0.056 | $0.016 |
| 30-turn chat, growing context, cached | $0.277 | $0.329 | $0.095 |
Source: CalculatorAI · calculatorai.app · Same model and assumptions as above
Summarise a 20-page PDF
- OpenAI
- $0.0040
- Anthropic
- $0.019
- $0.0065
Support reply, cached 3k system prompt
- OpenAI
- $0.0005
- Anthropic
- $0.0023
- $0.0010
Translate 1,000 words
- OpenAI
- $0.0022
- Anthropic
- $0.0095
- $0.0044
Coding-agent turn, 60k context, 90% cached
- OpenAI
- $0.0047
- Anthropic
- $0.021
- $0.0084
30-turn chat, growing context, cached
- OpenAI
- $0.028
- Anthropic
- $0.126
- $0.050
| Task | OpenAI | Anthropic | |
|---|---|---|---|
| Summarise a 20-page PDF | $0.0040 | $0.019 | $0.0065 |
| Support reply, cached 3k system prompt | $0.0005 | $0.0023 | $0.0010 |
| Translate 1,000 words | $0.0022 | $0.0095 | $0.0044 |
| Coding-agent turn, 60k context, 90% cached | $0.0047 | $0.021 | $0.0084 |
| 30-turn chat, growing context, cached | $0.028 | $0.126 | $0.050 |
Source: CalculatorAI · calculatorai.app · Same model and assumptions as above
The task shapes: the summary is 15,000 tokens in and 800 out, nothing cached. The support reply is a 3,000-token system prompt plus a 500-token question, with 85% of the input served from cache, and a 300-token answer. The translation is 1,500 tokens in and 1,600 out. The coding-agent turn is 60,000 tokens of context (90% cached) and 2,000 tokens written. The chat is thirty turns of 150-token questions and 400-token answers over an 800-token system prompt, with each turn re-sending everything before it and 85% of that prefix hitting the cache.
What the tables actually sayThree patterns
The tier gap dwarfs the provider gap. Within a tier, the spread from cheapest to dearest provider is about 3× at the top two tiers and 4–5× at the bottom. Across tiers, the same provider's flagship costs 10–20× its own cheap model: a summary is $0.076 on GPT-5.6 Sol and $0.004 on GPT-5.6 Luna. If you are choosing a provider to save money, you are looking at the wrong knob. The knob is which tier each task actually needs — the routing decision our breakdown of what AI actually costs calls the one that dominates every bill.
Gemini wins every like-for-like row. Google prices its Pro at OpenAI's and Anthropic's mid-tier level and its Flash at a third of theirs, so at the flagship and mid tiers Gemini is roughly 2–3.5× cheaper for identical token shapes. Whether the quality is identical is a question this article cannot answer and does not try to; it is a price comparison, and the honest way to use it is to run your own evaluation set on the two candidates and then let this table break the tie.
Anthropic is the most expensive at every tier, and the tokenizer is part of why. Opus 5 and Sonnet 5 are a dollar or two above the OpenAI equivalents per million tokens, and then bill 30% more tokens for the same text. Haiku 4.5, priced at $1 in and $5 out, is closer to a mid-tier price than to Luna's or Flash-Lite's, which is why the cheap-tier column flips to OpenAI. None of this is a verdict on the models; it is a reason to count tokens per model rather than per text, which the Token Counter does.
A month of one workload, mid-tier
3,000 support replies (the cached-system-prompt shape) plus 400 coding-agent turns, priced at September 2026 rates. A $20 subscription buys a chat window with its own limits, not an API budget; whether it absorbs this workload is a separate question, worked through for OpenAI and Anthropic in the linked comparisons.
Caching is the other half of the priceIt is not a provider feature
Look at the coding-agent turn again. Without a cache it is 60,000 fresh input tokens every step; with the 90% hit rate in the table, it is 6,000 fresh and 54,000 at a tenth of the price. On the mid-tier models that turns $0.144 into $0.047 on GPT-5.6 Terra, $0.182 into $0.056 on Sonnet 5, and $0.052 into $0.016 on Gemini 3.8 Flash — a 68–69% cut on all three, because all three price a cache read at the same tenth of base input.
That equality is the point. Caching is not something one provider offers and another does not; it is a property of your request shape. A task with a large fixed prefix — a system prompt, a set of tool definitions, a document you keep asking about — is cheap on every provider if the prefix is sent byte-identical each time, and expensive on every provider if it is not. Our MCP token-cost analysis shows the extreme case: tool definitions that are 99% of the input and 100% cacheable.
Put the fixed part first and keep it identical
Caches match on an exact prefix. A timestamp, a user name or a reordered tool list at the top of the prompt breaks the match for everything after it.
Price the output separately
Output is 5–6× the input rate everywhere. In the translation task, output is 84–86% of the bill on every flagship. Shorter answers save more than shorter prompts.
Route by task, not by habit
Summaries, classification, extraction and routine replies belong on the cheap tier. The flagship is for the 10% of requests that fail on the mid-tier in your own evaluation.
Count tokens per model
The same document is 30% more tokens on Claude 4.7+ than on the OpenAI tokenizer. Compare requests in dollars, not in words.
Re-check the rates
These prices were read on September 9, 2026. Providers reprice a few times a year; the calculator says out loud when its table is older than 90 days.
When the cheapest column is the wrong answerThree honest caveats
Quality is not in the table. A model that costs a third as much and fails 5% more often is not cheaper if a human has to catch the failures. The only defensible way to choose is to run a fixed set of your real tasks through two candidates and score the outputs; then price the winner and the runner-up here. Our Claude Code vs Cursor vs Codex comparison takes that approach for coding agents, where the cost of a wrong answer is a broken build rather than an awkward sentence.
Long contexts have their own tiers. Google raises its rate above a 200,000-token prompt (Gemini 3.1 Pro goes from $2 to $4 input, $12 to $18 output), and the table above stays under that line. A request that stuffs a whole repository into the prompt is a different price on Gemini than the same request just under the threshold — and, as with caching, that is a request-shape decision.
Subscriptions are a different product. ChatGPT Plus, Claude Pro and Google AI Pro each cost about $20 a month and sell a chat interface with usage limits, not tokens. For a person who sends requests by hand they are often cheaper than the API; for anything that runs a loop they usually are not. We priced both sides for ChatGPT Plus vs Pro vs API and Claude Pro vs Max vs API; the Subscription vs API calculator runs your own volume through the same comparison.
Where these numbers come from
Every rate is read from src/config/ai-models.ts, the price table shared by CalculatorAI's five AI-cost calculators, which records each provider's pricing page and the date it was last verified (September 9, 2026). The five task shapes are fixed token counts stated in the text; the chat task is simulated turn by turn with the context re-sent each time and 85% of the resent prefix served from cache. The Anthropic v2 tokenizer factor of 1.3× is Anthropic's own published approximation and is applied to Opus 5 and Sonnet 5 only; Gemini is counted 1:1 because Google publishes no equivalent, which if anything flatters Google slightly. No batch discounts, long-context surcharges, tool charges or retries are included, so every figure is a floor for a well-behaved request. The script is drafts/chatgpt-vs-claude-vs-gemini-cost-per-task-numbers.mts; change a token count or a hit rate and every table recomputes. Model quality, latency and rate limits were not measured and are not claimed.
Frequently asked questions
Which is cheapest: ChatGPT, Claude or Gemini API? At the same tier, Gemini. At September 2026 rates Gemini 3.1 Pro costs about a third of Claude Opus 5 and half of GPT-5.6 Sol on identical tasks, and Gemini 3.8 Flash is about a third of GPT-5.6 Terra or Claude Sonnet 5. The single cheapest model on the list is OpenAI's GPT-5.6 Luna.
Is Claude more expensive than ChatGPT? Per task, yes, at every tier: a dollar or two more per million tokens, and Claude 4.7 and later count roughly 30% more tokens for the same text. A 20-page summary is about $0.12 on Opus 5 against $0.08 on GPT-5.6 Sol; on the mid-tier the gap narrows to $0.049 against $0.040.
How much does one AI support reply cost? With a 3,000-token system prompt served mostly from cache and a 300-token answer: about half a cent on the mid-tier models (GPT-5.6 Terra $0.005, Sonnet 5 $0.006, Gemini 3.8 Flash $0.002) and a fraction of a cent on the cheap tier. Three thousand replies a month is $5–$18.
Does prompt caching work the same on all three? Effectively yes. OpenAI, Anthropic and Google all bill a cache read at roughly a tenth of base input, so a request whose input is 90% cached costs about a third of the uncached version on any of them. What differs is the mechanics — automatic prefix matching versus explicit cache controls — not the discount.
Should I just use the cheapest model for everything? Use the cheapest tier that passes your own evaluation for each task type. Most classification, extraction and routine replies do; most multi-step reasoning does not. A flagship for everything costs 10–20× the cheap tier for the majority of requests that never needed it.






