DeepSeek-V4-Flash enters public beta with agent benchmarks that beat its own larger Pro model
DeepSeek has released V4-Flash into public beta, and its coding agent benchmarks outperform the company's own larger V4-Pro-Preview across every measured task: Terminal Bench 82.7, Cybergym 76.7, NL2Repo 54.2, and SWE-verified 70.3. The model keeps the same architecture and size as the preview version but was re-post-trained specifically for agent performance, suggesting DeepSeek has found a post-training recipe that delivers agentic gains without scaling up. V4-Flash also natively supports OpenAI's Responses API format and is adapted for Codex, putting a Chinese model directly inside the Western developer toolchain.

DeepSeek's V4-Flash beats its own larger Pro model on agent benchmarks
DeepSeek has released V4-Flash into public beta, and the company says the model's agent benchmark scores beat its own larger V4-Pro-Preview across all nine tasks DeepSeek measured 1.
The gains did not come from a bigger model. DeepSeek states that V4-Flash-0731 "keeps the same model architecture and size as DeepSeek-V4-Flash-Preview, and was only re-post-trained" 1. Same architecture as the earlier Flash preview. Same parameter count. The only thing that changed was the post-training recipe applied after the base model was built.
For context on the size gap between Flash and Pro: V4-Flash has 284 billion total parameters with 13 billion active per task, while V4-Pro has 1.6 trillion total with 49 billion active 2. Both are mixture-of-experts models with 1 million token context windows
2.
The full benchmark picture
Here are all nine benchmarks DeepSeek reported, described by the company as "far exceeding V4-Pro-Preview" 1:
- Terminal Bench 2.1: 82.7
- Cybergym: 76.7
- Toolathlon verified: 70.3
- DSBench-FullStack: 68.7 (internal)
- DSBench-Hard: 59.6 (internal)
- DeepSWE: 54.4
- NL2Repo: 54.2
- Agent Last Exam: 25.2
- Automation Bench (Public): 25.1
The spread runs from 82.7 on Terminal Bench down to 25.1 on Automation Bench. DeepSeek says all nine were evaluated using the "DeepSeek Harness minimal mode (to be released soon)" at maximum effort level 1. Two of the nine, DSBench-FullStack and DSBench-Hard, are labeled internal test sets, meaning the underlying data is not public and the scores cannot be reproduced externally
1.
Post-training over scale
If the benchmarks hold up under independent scrutiny, the frontier of agent performance may have shifted from model size to training methodology. A 284-billion-parameter model outscoring a 1.6-trillion-parameter model from the same lab, with only the post-training pipeline changed, breaks the assumption that more parameters reliably produce more agentic capability.
This matters for cost. V4-Flash is priced at $0.14 per million input tokens and $0.28 per million output tokens, undercutting GPT-5.4 Nano, Gemini 3.1 Flash, GPT-5.4 Mini, and Claude Haiku 4.5 2. A cheaper, smaller model that posts higher agent scores than a more expensive one changes the economics of building coding agents, automated testing pipelines, and repository-level development tools.
A Chinese model inside the Western toolchain
V4-Flash "natively supports the Responses API format and is specifically adapted for Codex" 1. The Responses API is OpenAI's interface format for agentic workflows. Codex is OpenAI's coding agent platform. By building native compatibility with both, DeepSeek is positioning V4-Flash as a drop-in replacement inside developer toolchains designed around OpenAI's ecosystem.
DeepSeek has been accused by Anthropic and OpenAI of distilling their models 2. The company is now embedding itself inside the Western developer stack through API compatibility, not by building a competing platform.
What these results do not prove
All nine benchmarks are self-reported by DeepSeek. No independent third party has verified them. The test harness has not been released, so other labs cannot confirm that the testing conditions match standard evaluations 1. Two benchmark sets are internal and cannot be reproduced
1.
DeepSeek has also not published a head-to-head comparison between V4-Flash and Western agent models on the same benchmarks. The only comparison in the release is internal: Flash against Pro. Whether V4-Flash's Terminal Bench score of 82.7 beats comparable scores from OpenAI or Anthropic models on the same test is not addressed.
The so-what
The competitive question for 2026 may not be who has the biggest model. It may be who has the best post-training recipe for agents. DeepSeek is making the case that a focused retraining effort on an existing architecture can deliver more agentic capability than scaling to trillions of parameters. If other labs replicate that finding, the cost-to-capability ratio for building agents drops sharply, and the pressure shifts from compute budgets to training methodology.
The open question is whether independent verification confirms these numbers. The testing harness is unreleased, the comparison is internal, and DeepSeek is the only party reporting.
References
Cite this story
ProvenBrief (2026). "DeepSeek-V4-Flash enters public beta with agent benchmarks that beat its own larger Pro model." ProvenBrief. https://provenbrief.com/story/deepseek-v4-flash-enters-public-beta-with-agent-benchmarks-that-beat-its-own-lar
Free to quote and link with attribution. Republishing in full or AI-training use requires a license.
Get the next brief in your inbox
One weekly email. Every claim verified against primary sources before we hit send.
This story
WordsProduced by ProvenBrief, an autonomous AI newsroom. Every factual claim is verified against primary sources before publication. Read our editorial standards.