DeepSeek V4 Flash is available in Agent Planners: a 1M-context, ultra-low-cost model
DeepSeek V4 Flash is a selectable model across every agent in Agent Planners: a 1,000,000-token context window at $0.14 per 1M input tokens and $0.28 per 1M output tokens on the live AI Gateway — among the cheapest per-token rates of any model in the catalog, on a context window that matches models many multiples of its price. It's not a new addition rushed out for a headline; it's been a curated, correctly-billed option, and this is the real pricing and the real fit.
The current numbers
| DeepSeek V4 Flash | DeepSeek V4 Pro | GPT-5.6 Luna (default) | |
|---|---|---|---|
| Context window | 1,000,000 tokens | 1,000,000 tokens | 1,050,000 tokens |
| Price per 1M tokens | $0.14 input / $0.28 output | $0.435 input / $0.87 output | $0.20 input / $1.20 output |
| Credit rate / 1K tokens | 1 in · 4 out | 1 in · 4 out (strong-tier rule) | 1 in · 6 out |
| Best suited for | High-volume fast-tier steps needing a big context window | Long-context research and synthesis | Quality-critical planning and optimization |
Where it fits among six agents
Agent Planners runs six specialized agents — orchestrator, analytics, optimization, creative, safety and research — and lets you pick a model per agent rather than one model for the whole system. DeepSeek V4 Flash's combination of a large context window at a very low per-token cost makes it a strong fit anywhere you want headroom to read a lot of account or report data cheaply: research and analytics in particular, where the value is in surfacing the right data, not in the deepest possible reasoning chain.
For where a premium or strong-tier model earns its price instead — the orchestrator on a genuinely hard multi-step goal, for instance — see the full decision framework for choosing a model per agent.
New models keep surfacing without a product update
Agent Planners' model picker and public pricing pages sync from the live AI Gateway catalog, not a hand-maintained list alone — so a fresh dated model snapshot the moment it appears on the gateway shows up on the AI Model Explorer and the model pricing & credits ranking automatically, correctly priced from the real gateway rate. Curating a model with a clean label (as we've done here for DeepSeek V4 Flash) is a naming convenience on top of that — the model itself doesn't wait on it to be usable.
How billing stays accurate for a model like this
Every model's credit rate is metered in order of authority: the live gateway price first, on every call; a daily-refreshed persisted snapshot if that call is briefly unreachable; and a static fallback table as the last resort. DeepSeek V4 Flash has an explicit fallback rule (1 credit per 1K input tokens, 4 per 1K output) rather than relying on a generic catch-all — the same audit discipline that caught a real 30x billing gap for GPT-5.6 Luna's fallback rule, described in the Luna default-model story.
Frequently asked questions
- Is DeepSeek V4 Flash available in Agent Planners?
- Yes — it's a curated, selectable model on every agent, with correctly-billed pricing: $0.14 per 1M input tokens and $0.28 per 1M output tokens on a 1,000,000-token context window, verified against the live AI Gateway catalog.
- How does DeepSeek V4 Flash compare to DeepSeek V4 Pro?
- Flash is roughly 3x cheaper per token than Pro on the same 1,000,000-token context window. Pro is the better fit for the hardest long-context synthesis; Flash is the better fit for high-volume steps where you want a large context window without paying Pro's rate.
- What's the best agent to run DeepSeek V4 Flash on?
- Research and analytics are strong fits — both benefit from a large context window more than from the deepest possible reasoning chain, and Flash's low per-token cost makes reading a lot of account or report data cheap.
- Do I need to wait for a product update to use a brand-new model?
- Not for basic selectability — Agent Planners' model picker and public pricing pages sync from the live AI Gateway catalog directly, so newly released models can surface automatically. A curated label and dedicated billing rule (like DeepSeek V4 Flash has now) is a polish step on top, not a prerequisite.
- Is DeepSeek V4 Flash's pricing likely to be accurate going forward?
- Billing checks the live gateway price on every call first, falling back to a daily-refreshed snapshot and then a static table only if that call is briefly unreachable — so pricing stays correct even as gateway rates change, rather than drifting against a number typed in once.