August 3, 2026 · 8 min read
Product

AI model pricing comparison (2026): what GPT, Claude, Gemini and DeepSeek actually cost

TL;DR

Per-token prices across current frontier and budget models span roughly 50x on input and over 100x on output — and the cheapest model isn't automatically the wrong choice, or the most expensive the right one. This is a snapshot of real, live pricing (verified against the AI Gateway catalog, not vendor spec sheets) across seven models spanning ultra-cheap to premium, and a plain framework for what the spread actually means when you're picking a model for a specific job rather than shopping for the whole market.

The current numbers, per 1M tokens

ModelInputOutputContext window
DeepSeek V4 Flash$0.20$0.401,000,000
GPT-5.6 Luna$0.20$1.201,050,000
DeepSeek V4 Pro$0.44$0.871,000,000
Gemini 3.5 Flash$1.50$9.001,000,000
Claude Sonnet 4.6$3.00$15.001,000,000
GPT-5.5$5.00$30.001,000,000
Claude Opus 5$5.00$25.001,000,000
Claude Fable 5$10.00$50.001,000,000
Verified live against the AI Gateway catalog on August 3, 2026 — base rates, before regional or long-context tier adjustments some providers apply above ~272K tokens in a single call. Provider pricing changes; treat this as a snapshot, not a permanent number.

The pattern underneath the numbers

Context window size barely correlates with price here — every model above sits at roughly 1,000,000 tokens regardless of whether it costs $0.20 or $10 per million input tokens. What actually drives the spread is reasoning depth and output quality on hard, judgment-heavy tasks: the cheapest models are excellent at fast, structured, low-ambiguity work, and the premium tier earns its price specifically on multi-step planning and creative output where a wrong call is expensive to have missed.

Output almost always costs more than input, and the ratio varies a lot by provider — Claude's models run roughly 5x output-vs-input, while GPT-5.5's ratio is 6x and DeepSeek's is closer to 2x. This matters operationally: a task that generates a lot of output (a long report, extensive ad copy) is more exposed to that ratio than a task that mostly reads and summarizes.

What this means for picking a model per agent

The 50x spread is the entire point of Agent Planners running six specialist agents on independently chosen models rather than one model for everything: a safety agent doing structured yes/no checks gets zero benefit from a $10-per-million model, and an orchestrator planning a genuinely hard multi-step goal is the one place that price is usually worth paying. See the full decision framework for choosing a model per agent for how to map this table to your own six agents rather than picking one model for the whole system.

Where to check current pricing yourself

This table is a snapshot; for pricing that's always current, two tools pull directly from the live AI Gateway catalog rather than a maintained document: the AI model pricing & credits ranking ranks every selectable model cheapest-to-priciest by credit rate, and the AI Model Explorer lets you search, filter and compare models side by side interactively.

Frequently asked questions

Which AI model is cheapest per token right now?
DeepSeek V4 Flash and GPT-5.6 Luna are the cheapest of the models compared here at $0.20 per 1M input tokens, though their output rates differ ($0.40 vs $1.20 per 1M) — Flash is cheaper end-to-end for output-heavy tasks.
Why is Claude Fable 5 so much more expensive than Claude Sonnet 4.6?
Fable 5 is Anthropic's premium tier, priced roughly 3-3.3x Sonnet 4.6's rate ($10/$50 vs $3/$15 per 1M). The premium tier is priced for the hardest judgment-heavy or creative work, not routine tasks.
Does a bigger context window cost more?
Not directly in this comparison — every model listed sits at roughly 1,000,000 tokens of context regardless of price. The price spread tracks reasoning/output quality, not context size.
Is the cheapest model always the wrong choice for important work?
No — cheap, fast models are often the right choice for structured checks and routine data pulls. The premium tier earns its price specifically on planning and creative work where the difference in output quality is large and visible.
How current is this pricing table?
Verified live against the AI Gateway catalog on August 3, 2026. For pricing that's always current rather than a dated snapshot, use the model pricing ranking or AI Model Explorer, both of which pull the live catalog directly.
Run your ad ops with an agent you can trust

Start free — 2,500 credits a month, no credit card. Every write waits for your approval.

Keep reading