DeepSeek V4 Flash Benchmarks
The official 0731 release was only re-post-trained — same architecture, same size — yet agent benchmarks now far exceed V4 Pro Preview.
Agentic Benchmarks
| Benchmark | V4 Flash Score | Note |
|---|---|---|
| Terminal-Bench 2.0 | 49.1% | Exceeds V4 Pro Preview |
| MCP Atlas | 64% | Tool-use orchestration |
| Toolathlon | 40.7% | Multi-tool chaining |
| Claw-Eval | 57.8% | Agent reliability |
| Gert Labs | 54.35% | Real-world agent tasks |
Coding Benchmarks
| Benchmark | V4 Flash Score | Note |
|---|---|---|
| SWE-bench Verified | Top tier | Open-weight leader |
| LiveCodeBench | Competitive | vs models 4x the size |
| Aider Polyglot | Strong | Multi-language coding |
| HumanEval+ | High | Code generation accuracy |
| Terminal-Bench Coding | Leading | CLI-based coding tasks |
Reasoning Benchmarks
| Benchmark | V4 Flash Score | Note |
|---|---|---|
| GPQA Diamond | Competitive | Graduate-level QA |
| MATH-500 | Strong | Mathematical reasoning |
What the Numbers Mean
The headline: a 13B-active model now matches or exceeds the agent performance of its 49B-active sibling. DeepSeek says the architecture is identical to the preview — the improvement comes entirely from post-training. That means V4 Pro's official release will likely close the gap, but for now V4 Flash is the strongest agent model per dollar in the open-weight space.
Caveat: multilingual and multimodal benchmarks are not measured. V4 Flash is a text-only model optimized for agent, coding, and reasoning tasks. If you need vision or audio, look at DeepSeek's separate VL line.
Benchmark FAQ
Why does V4 Flash beat V4 Pro on agent tasks?
The 0731 release applies improved post-training specifically targeting agent capabilities. DeepSeek's changelog states the architecture is unchanged — the gains come entirely from better training, suggesting V4 Pro's official release will likely improve too.
What is Terminal-Bench 2.0?
A benchmark that evaluates AI models on real terminal-based coding and system administration tasks. V4 Flash scores 49.1%, which exceeds V4 Pro Preview — notable because Pro has nearly 4x the active parameters.
What is MCP Atlas?
MCP Atlas tests a model's ability to use tools through the Model Context Protocol. V4 Flash scores 64%, indicating strong capability in multi-tool orchestration — a critical skill for AI agents.
Are these benchmarks verified?
Agentic benchmarks (5 of 5) and coding benchmarks (5 of 5) are verified from published sources. Reasoning has 2 verified scores. Multilingual and multimodal categories are not measured — It is text-only.