August 1, 2026 · 6 min read
Product

DeepSeek V4 Flash is available in Agent Planners: a 1M-context, ultra-low-cost model

TL;DR

DeepSeek V4 Flash is a selectable model across every agent in Agent Planners: a 1,000,000-token context window at $0.14 per 1M input tokens and $0.28 per 1M output tokens on the live AI Gateway — among the cheapest per-token rates of any model in the catalog, on a context window that matches models many multiples of its price. It's not a new addition rushed out for a headline; it's been a curated, correctly-billed option, and this is the real pricing and the real fit.

The current numbers

DeepSeek V4 FlashDeepSeek V4 ProGPT-5.6 Luna (default)
Context window1,000,000 tokens1,000,000 tokens1,050,000 tokens
Price per 1M tokens$0.14 input / $0.28 output$0.435 input / $0.87 output$0.20 input / $1.20 output
Credit rate / 1K tokens1 in · 4 out1 in · 4 out (strong-tier rule)1 in · 6 out
Best suited forHigh-volume fast-tier steps needing a big context windowLong-context research and synthesisQuality-critical planning and optimization
Prices verified directly against the live AI Gateway catalog, not a spec sheet — this is what the model actually costs to call right now, and it's what Agent Planners' billing meters against by default (see the order of authority below).

Where it fits among six agents

Agent Planners runs six specialized agents — orchestrator, analytics, optimization, creative, safety and research — and lets you pick a model per agent rather than one model for the whole system. DeepSeek V4 Flash's combination of a large context window at a very low per-token cost makes it a strong fit anywhere you want headroom to read a lot of account or report data cheaply: research and analytics in particular, where the value is in surfacing the right data, not in the deepest possible reasoning chain.

For where a premium or strong-tier model earns its price instead — the orchestrator on a genuinely hard multi-step goal, for instance — see the full decision framework for choosing a model per agent.

New models keep surfacing without a product update

Agent Planners' model picker and public pricing pages sync from the live AI Gateway catalog, not a hand-maintained list alone — so a fresh dated model snapshot the moment it appears on the gateway shows up on the AI Model Explorer and the model pricing & credits ranking automatically, correctly priced from the real gateway rate. Curating a model with a clean label (as we've done here for DeepSeek V4 Flash) is a naming convenience on top of that — the model itself doesn't wait on it to be usable.

How billing stays accurate for a model like this

Every model's credit rate is metered in order of authority: the live gateway price first, on every call; a daily-refreshed persisted snapshot if that call is briefly unreachable; and a static fallback table as the last resort. DeepSeek V4 Flash has an explicit fallback rule (1 credit per 1K input tokens, 4 per 1K output) rather than relying on a generic catch-all — the same audit discipline that caught a real 30x billing gap for GPT-5.6 Luna's fallback rule, described in the Luna default-model story.

Frequently asked questions

Is DeepSeek V4 Flash available in Agent Planners?
Yes — it's a curated, selectable model on every agent, with correctly-billed pricing: $0.14 per 1M input tokens and $0.28 per 1M output tokens on a 1,000,000-token context window, verified against the live AI Gateway catalog.
How does DeepSeek V4 Flash compare to DeepSeek V4 Pro?
Flash is roughly 3x cheaper per token than Pro on the same 1,000,000-token context window. Pro is the better fit for the hardest long-context synthesis; Flash is the better fit for high-volume steps where you want a large context window without paying Pro's rate.
What's the best agent to run DeepSeek V4 Flash on?
Research and analytics are strong fits — both benefit from a large context window more than from the deepest possible reasoning chain, and Flash's low per-token cost makes reading a lot of account or report data cheap.
Do I need to wait for a product update to use a brand-new model?
Not for basic selectability — Agent Planners' model picker and public pricing pages sync from the live AI Gateway catalog directly, so newly released models can surface automatically. A curated label and dedicated billing rule (like DeepSeek V4 Flash has now) is a polish step on top, not a prerequisite.
Is DeepSeek V4 Flash's pricing likely to be accurate going forward?
Billing checks the live gateway price on every call first, falling back to a daily-refreshed snapshot and then a static table only if that call is briefly unreachable — so pricing stays correct even as gateway rates change, rather than drifting against a number typed in once.
Run your ad ops with an agent you can trust

Start free — 2,500 credits a month, no credit card. Every write waits for your approval.

Keep reading