"How much will our AI agent cost?" has no single answer, because "agent" covers everything from a two-step email sorter to a coding assistant that makes twenty-five model calls to fix one bug. What it does have is a formula, and the formula is not "tokens per request × requests".
We priced five agents that companies actually run — email triage, a support bot, a knowledge-base assistant, a research agent and a coding agent — at realistic monthly volumes, on the API rates each provider published as of October 1, 2026. With prompt caching switched on, the monthly bills ranged from $6 to $2,222. The same agents with caching off cost up to 4.7× more, and the model you pick moves the bill by another 5–20×.
The answerFive agents, priced per month
Each row is one workload at its monthly volume, with the whole conversation cached (and the cache-write fee included where the provider charges one). The three model columns go from the cheapest model that is plausible for the job to a premium one.
Email triage (2 calls)
- Volume
- 20,000 emails
- Lowest-cost fit
- $16 · GPT-6 Luna
- Middle
- $161 · Claude Haiku 4.5
- Premium
- $321 · Claude Sonnet 5.5
Support bot (4 calls)
- Volume
- 3,000 tickets
- Lowest-cost fit
- $6 · GPT-6 Luna
- Middle
- $43 · Gemini 3.8 Flash
- Premium
- $129 · Claude Sonnet 5.5
Knowledge-base / RAG (2 calls)
- Volume
- 10,000 questions
- Lowest-cost fit
- $18 · GPT-6 Luna
- Middle
- $117 · Gemini 3.8 Flash
- Premium
- $361 · Claude Sonnet 5.5
Research agent (12 calls)
- Volume
- 300 reports
- Lowest-cost fit
- $38 · Gemini 3.8 Flash
- Middle
- $114 · Claude Sonnet 5.5
- Premium
- $202 · Claude Opus 5.5
Coding agent (25 calls)
- Volume
- 400 tasks
- Lowest-cost fit
- $444 · Claude Sonnet 5.5
- Middle
- $740 · Claude Opus 5.5
- Premium
- $2,222 · GPT-6 Astra
| Agent | Volume | Lowest-cost fit | Middle | Premium |
|---|---|---|---|---|
| Email triage (2 calls) | 20,000 emails | $16 · GPT-6 Luna | $161 · Claude Haiku 4.5 | $321 · Claude Sonnet 5.5 |
| Support bot (4 calls) | 3,000 tickets | $6 · GPT-6 Luna | $43 · Gemini 3.8 Flash | $129 · Claude Sonnet 5.5 |
| Knowledge-base / RAG (2 calls) | 10,000 questions | $18 · GPT-6 Luna | $117 · Gemini 3.8 Flash | $361 · Claude Sonnet 5.5 |
| Research agent (12 calls) | 300 reports | $38 · Gemini 3.8 Flash | $114 · Claude Sonnet 5.5 | $202 · Claude Opus 5.5 |
| Coding agent (25 calls) | 400 tasks | $444 · Claude Sonnet 5.5 | $740 · Claude Opus 5.5 | $2,222 · GPT-6 Astra |
Source: CalculatorAI · calculatorai.app · CalculatorAI agent cost engine · platform.claude.com · developers.openai.com · ai.google.dev
Two things stand out before any detail. The high-volume agents (triage, support, RAG) are cheap per task and expensive only because of volume. The long agents (research, coding) are the opposite: a few hundred tasks a month, but each task is so large that they dominate any combined bill.
The loopWhy an agent bill is not a chatbot bill
A chatbot reply is one request. An agent works in a loop — think, call a tool, read the result, think again — and every step is a new API request that resends everything before it: the system prompt, the tool definitions, the task, and every earlier step's output and tool result.
So the requests are not the same size. In our coding agent the first call sends 14,000 tokens; the twenty-fifth sends 146,000, because it carries twenty-four steps of history. Across the task that adds up to 2.0 million input tokens against 37,500 tokens of output. The model writes about 2% of what it reads.
calls × (system + task) + (output + tool result per step) × calls × (calls − 1) ÷ 2(input tokens × input rate + output tokens × output rate) × tasks per monthThe second half of the first formula is the part people miss, and it grows with the square of the number of steps. Priced as "tokens per call × calls", our coding agent comes out at $1.08 a task on Claude Sonnet 5.5. Priced as a loop, it is $4.38 before caching — 4.1× more. Our earlier guide to what AI actually costs shows the same effect on a twelve-step research loop: seventeen times one step, not twelve.
Five workloadsWhat each agent does, and what it costs
The token shapes below are assumptions, chosen to be typical rather than flattering. Yours will differ; the point is which inputs drive the bill.
Email triage — 20,000 emails a month
Two calls per email: read and classify, then draft a routing note. 2,500 tokens of instructions and labels, a 1,500-token email, 200 tokens written per step.
On GPT-6 Luna that is $0.0008 per email and $16 a month. Claude Haiku 4.5 costs $161 and Claude Sonnet 5.5 $321 for the same work. Sorting email is the kind of job where the cheapest tier is often good enough — which is exactly what to test before paying 20× more.
Support bot — 3,000 tickets a month
Four calls per ticket: understand the question, look up the order, check the policy, reply. A 6,000-token system prompt (tone, refund policy, tool definitions), 800 tokens back from each tool call.
$6 a month on GPT-6 Luna, $43 on Gemini 3.8 Flash, $129 on Claude Sonnet 5.5 — 4.3 cents a ticket on the most expensive of the three. Caching matters more here than on triage (52% off versus 34%), because the 6,000-token prompt is resent on every one of the four calls.
Knowledge-base (RAG) assistant — 10,000 questions a month
Two calls: answer from 6,000 tokens of retrieved documentation, then check the answer against it. The retrieved text is different for every question, which is what caps the saving.
$18 on GPT-6 Luna, $117 on Gemini 3.8 Flash, $361 on Claude Sonnet 5.5 and $381 on GPT-5.6 Terra. A short RAG loop spends most of its money on fresh input that no cache can help with — the same reason our prompt caching guide found caching worth only 15–17% on single-shot RAG requests.
Research agent — 300 reports a month
Twelve calls: search, read a 6,000-token page, take notes, repeat, then write. The last call carries 79,300 tokens.
$38 on Gemini 3.8 Flash, $114 on Claude Sonnet 5.5, $202 on Claude Opus 5.5 — 67 cents a report on Opus. Without caching, the Opus bill would be $661.
Coding agent — 400 tasks a month
Twenty-five calls: read files, edit, run tests, read the output, fix, repeat. A 12,000-token system prompt with tool definitions, 4,000 tokens back from each tool, 1,500 tokens written per step. Four hundred tasks is roughly one developer running the agent twenty times a working day.
$444 on Claude Sonnet 5.5, $474 on GPT-5.6 Terra, $740 on Claude Opus 5.5 and $2,222 on GPT-6 Astra. Per task: $1.11, $1.19, $1.85 and $5.55. Whether a subscription plan beats these API figures for a single developer is a separate question, worked through in Claude Code vs Cursor vs Codex.
CachingThe lever that decides whether a long agent is viable
The AI Agent Cost Calculator prices three caching modes, and the gap between the last two is the most important number in this article.
None
- Per task
- $4.38
- Per month
- $1,750
- Saving
- —
System prompt and tools only
- Per task
- $3.86
- Per month
- $1,543
- Saving
- 12%
Whole conversation (reads only)
- Per task
- $1.04
- Per month
- $415
- Saving
- 76%
Whole conversation, write fee included
- Per task
- $1.11
- Per month
- $444
- Saving
- 75%
| Caching | Per task | Per month | Saving |
|---|---|---|---|
| None | $4.38 | $1,750 | — |
| System prompt and tools only | $3.86 | $1,543 | 12% |
| Whole conversation (reads only) | $1.04 | $415 | 76% |
| Whole conversation, write fee included | $1.11 | $444 | 75% |
Source: CalculatorAI · calculatorai.app · CalculatorAI agent cost engine · platform.claude.com
Caching only the system prompt — the common hand-rolled setup — barely helps a long agent, because by step ten the system prompt is a small share of what is being resent. The history is the expensive part, and only caching the whole conversation, so that each step pays full price for just the tokens the previous step added, takes it off the bill. On Opus 5.5, whose cache reads cost 0.05× input rather than the usual 0.1×, the same switch cuts the bill by 79%, from $3,500 to $740.
Caching also narrows the gap between models. Uncached, Opus 5.5 costs exactly twice Sonnet 5.5 on the coding agent; cached, it costs 1.67×, because Opus's cache discount is deeper. The cheaper the reads, the more an expensive model's price is really its output price.
Model choiceSame work, 5× to 20× apart
Within each workload, the cheapest and the most expensive plausible model differ by 5–20× on the cached bill: about 20× on the three high-volume agents, about 5× on research and coding. That spread is larger than anything caching or batching can do, which makes model choice the first decision, not the last.
$1,369 a month
All five agents on Claude Sonnet 5.5, conversation cached, write fee included: $321 + $129 + $361 + $114 + $444.
$698 a month
GPT-6 Luna for triage and support, Gemini 3.8 Flash for RAG, Sonnet 5.5 kept for research and coding. Half the bill, if the cheap tiers pass your quality test.
That "if" carries the whole decision. A support reply that is wrong costs a customer; a misrouted email costs a few seconds. Run a few hundred real tasks through the cheaper model and read them before switching — the rates themselves are compared provider by provider in ChatGPT vs Claude vs Gemini: cost per real task.
Four more inputsWhat else moves the monthly number
Retries. A step that fails and is re-run costs again. At a 5% retry rate the support bot on Sonnet 5.5 goes from $115 to $120 a month (reads only) — small, but it scales every figure in this article by the same factor.
Batch processing. Anthropic and OpenAI price asynchronous batch requests at half the standard rate. Research reports that nobody needs within the minute are a natural fit: the research agent on Sonnet 5.5 falls from $102 to $51 a month (reads only) when batched.
Long-context surcharges. Gemini 3.1 Pro, Grok and some other models raise the rate on any request above 200,000 tokens. None of our five agents crosses it — the coding agent peaks at 146,000 — but add ten more steps to it and the last few calls would. The calculator applies the surcharge call by call, only to the calls that cross the threshold.
Tool definitions. Every tool and MCP server you connect sits in the system prompt and is resent on every step. Our breakdown of MCP server token costs shows how fast that grows.
Price your ownFour numbers to measure first
Calls per task
Count them from logs, not from the design doc. Agents loop more than planned, and the cost grows with the square of this number.
Tool result size
What a search, file read or database lookup returns, in tokens. It is added to the history on every step that follows.
System prompt + tool definitions
Everything resent on every call. Measure it once; it is usually bigger than people think.
Tasks per month
The multiplier. A cheap task at high volume and an expensive task at low volume can land on the same bill.
Put those four into the AI Agent Cost Calculator, pick the caching mode your provider actually gives you, and compare two or three models. For a single request rather than a loop, the AI Token Cost Calculator is the faster tool.
MethodWhere these numbers come from
Every figure was computed with the same agent-cost engine the calculator uses, on the price catalogue we re-verified against each provider's own pricing page on October 1, 2026: Anthropic, OpenAI and Google.
- Token shapes are assumptions, stated with each workload: calls per task, system prompt, task input, output and tool result per step. Real agents vary from task to task; these are typical, not measured from one product.
- Each step resends the whole history, the standard behaviour of a tool-calling agent loop.
- Nothing is cached on the first call of a task. At high volume a shared system prompt would usually still be warm in the cache from the previous task, so the short, high-volume agents (triage, support, RAG) are slightly overstated.
- The cache-write fee (+0.25× input on every full-price input token) was added for Anthropic and OpenAI GPT-5.6/GPT-6 in the cached rows. Google charges no write fee.
- No retries and no batch discount unless the text says so, so most figures are floors.
- The same token counts are used on every model. Tokenizers differ, so the same text can be a few percent more or fewer tokens on another provider.
- DeepSeek is left out: the model listed on its pricing page did not match the catalogue's model id on the day we checked, and we would rather omit a price than guess one.
FAQAI agent costs
How much does an AI agent cost per month?
In our five modeled workloads, between $6 a month (a support bot on GPT-6 Luna, 3,000 tickets) and $2,222 (a coding agent on GPT-6 Astra, 400 tasks), with prompt caching on. The main drivers are calls per task, how much each step adds to the conversation, the model, and monthly volume.
Why is an AI agent more expensive than a chatbot?
Because one agent task is many API calls, and each call resends the entire conversation so far. Input tokens grow with the square of the number of steps, so a 25-step task costs far more than 25 single requests.
Does prompt caching really cut agent costs that much?
Caching the whole conversation cut our research and coding agents' bills by 66–79%, including the write fee where it applies. Caching only the system prompt cut the same bills by 7–12%, because the growing history is the expensive part.
Which model is cheapest for an AI agent?
For simple, high-volume jobs like email triage, a budget tier such as GPT-6 Luna costs about a twentieth of Claude Sonnet 5.5. For long reasoning-heavy loops like coding, the cheapest model that reliably finishes the task is usually cheaper overall than a budget model that needs more steps or retries.
Are subscriptions cheaper than the API for coding agents?
For one developer, often yes — plans bundle a usage allowance. For automated agents running without a person, the API is the only option and these per-task figures apply.






