Input tokens, not output, drive AI agent costs — measured on 49 models
Input tokens, not output, are most of what an AI agent's model calls cost: on the median production task the agent read 32 tokens for every one it wrote, and output was under half of the model cost on all 49 models we priced. The usual advice to compare models by output price picks the wrong model for agent work.
Why the chat-era rule breaks for agents
A chat reply is a short prompt and a long answer, so output price dominates and comparing models by output price made sense. An agent works the other way round. Every step re-reads its instructions, the tool catalogue, the results of earlier steps and whatever data it pulled — and then writes a short tool call or a paragraph.
The median task in our measurement read 78,098 input tokens and wrote 2,450. A large task (90th percentile) read 87,357 and wrote 17,422 — still five to one.
Where the money goes, model by model
| Model | Input $/1M | Output $/1M | Cache read $/1M | Fresh input | Cache reads | Output | Median task (credits) |
|---|---|---|---|---|---|---|---|
| GPT-6 Luna | 0.10 | 0.50 | 0.010 | 58% | 13% | 29% | 29 |
| DeepSeek V4 Flash | 0.13 | 0.26 | 0.028 | 60% | 28% | 12% | 31 |
| Gemini 2.5 Flash | 0.30 | 2.50 | 0.030 | 48% | 11% | 41% | 101 |
| Llama 3.3 70B | 0.72 | 0.72 | 0.072 | 76% | 17% | 8% | 155 |
| DeepSeek V4 Pro | 0.66 | 1.98 | 0.022 | 73% | 5% | 22% | 163 |
| Gemini 3.8 Flash | 0.75 | 3.75 | 0.075 | 58% | 13% | 29% | 210 |
| Claude Haiku 4.5 | 1.00 | 5.00 | 0.100 | 58% | 13% | 29% | 280 |
| Grok 4.3 | 1.25 | 2.50 | 0.200 | 64% | 23% | 13% | 288 |
| Claude Sonnet 5 | 2.00 | 10.00 | 0.200 | 58% | 13% | 29% | 559 |
| Qwen3.7 Max | 2.50 | 7.50 | 0.500 | 57% | 25% | 17% | 617 |
| Claude Opus 5 | 5.00 | 25.00 | 0.500 | 58% | 13% | 29% | 1,397 |
| GPT-5.5 | 5.00 | 30.00 | 0.500 | 55% | 12% | 33% | 1,479 |
Three things the table shows
- Output is under half the bill on all 49 models for the median task — about 29% on the median model and as little as 8% on Llama 3.3 70B. Only on a large, output-heavy task does output reach half, on 14 of the 49.
- The cache-read price is a real line item. Where a model's cache read is a tenth of its input price, cache reads are about 13% of the cost; where it is a sixth to a fifth of the input price (Grok 4.3, Qwen3.7 Max, DeepSeek V4 Flash) they are 23–28%. DeepSeek V4 Pro's very cheap cache read brings it to 5%.
- Same ratio, different price. GPT-6 Luna, Claude Sonnet 5 and Claude Opus 5 have the same price shape, so the same split — but the same task costs 29, 559 and 1,397 credits. The shape tells you where to save; the level tells you how much.
What to do with it
- Compare models on input price first — and on cache-read price, because most of an agent's input is the same every step.
- Keep the stable part of the prompt stable. Instructions and tool definitions at the start, unchanged between calls, are what a cache can serve. A timestamp in the system prompt breaks that on every call.
- Read less. The cheapest token is one the agent never loads: select the accounts and dates the question needs, and let a step return a summary rather than every row.
- List fewer tools. A tool catalogue is input on every call. On our MCP server, listing tools on demand took a fully connected key from 171,665 input tokens a turn to about 4,800 — the write-up.
Frequently asked questions
- Do output tokens dominate the cost of an AI agent?
- No. Measured on production agent tasks, the median task read 32 input tokens for every output token, and output was under half of the model cost on all 49 models priced — about 29% on the median model. Output dominates chat replies, not agent work.
- How do I choose a cost-efficient model for an AI agent?
- Compare input price and cache-read price before output price, because an agent re-reads its instructions, tools and earlier results on every step. Then check the cost of a whole task, not a token: the same measured task ranged from 21 to 8,872 credits across 49 models.
- Does prompt caching matter for AI agents?
- Yes. In our measurement 43% of input tokens came from the prompt cache. Cache reads were 5–32% of a task's model cost depending on how a model prices them, so keeping the instructions and tool list unchanged between calls saves real money.