GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
← Back to feed

DeepSeek V4-Flash 0731 Hits 82.7 on Terminal Bench 2.1

DeepSeek has moved V4-Flash into public beta with a July 31 API update focused almost entirely on agentic coding and tool-use performance.

The important detail is what did not change. DeepSeek-V4-Flash-0731 keeps the same architecture and size as the preview release. The lift comes from re-post-training and a harness configuration aimed at code-agent workloads, not a larger base model.

The New Scores

BenchmarkDeepSeek V4-Flash 0731
Terminal Bench 2.182.7
NL2Repo54.2
Cybergym76.7
DeepSWE54.4
Toolathlon verified70.3
Agent Last Exam25.2
Automation Bench Public25.1
DSBench-FullStack68.7
DSBench-Hard59.6

Terminal Bench 2.1 at 82.7 is the headline because it puts V4-Flash in the same public range as the strongest coding-agent stacks, while preserving the model line DeepSeek positioned as the smaller, cheaper member of the V4 family.

The supporting numbers matter too. Toolathlon verified at 70.3 and Cybergym at 76.7 point to tool orchestration and security-adjacent task execution, not just patch generation. DeepSWE at 54.4 is less dominant, but it gives the result shape: V4-Flash is being tuned as a broad agent model rather than a single-benchmark specialist.

Codex Compatibility Is the Product Move

The release adds native Responses API support and a configuration specifically adapted for Codex-style workflows. That is not a small implementation note. The agent market has started to standardise around OpenAI-compatible APIs, long-running coding harnesses, and tool calls that look more like operating-system actions than chat completions.

DeepSeek is meeting that market where it is. V4-Flash can be called by setting the model name to deepseek-v4-flash, while the API surface stays stable. For teams already running OpenAI-style agents, the migration burden is deliberately low.

That matters because agent quality is increasingly a stack property. The base model, harness, tool schema, context policy, retry loop, and execution sandbox all contribute to the final result. DeepSeek is not only publishing a model score. It is trying to make V4-Flash a drop-in candidate for the same harness layer that routes work to GPT, Claude, Gemini, and open-weight coding models.

The Harness Caveat

The benchmark run used DeepSeek Harness minimal mode, max effort, top-p 0.95, and temperature 1.0. DeepSeek says the harness will be released later.

That caveat is material. Terminal Bench results are extremely sensitive to scaffold design, retry policy, shell affordances, and context handling. A raw model score and a production-agent score are not the same thing. Until the harness is public, the cleanest reading is that V4-Flash-0731 is a strong DeepSeek-stack result, not a fully independent apples-to-apples model result.

Still, the direction is clear. DeepSeek did not wait for V4-Pro to ship the agentic jump. It pushed the cheaper Flash line first.

What Comes Next

DeepSeek says V4-Pro API is unchanged for now and that the official V4-Pro release will follow soon. That sets up the next comparison: whether Pro brings a broader capability gain, or whether Flash becomes the better agentic price-performance product because it received the Codex-oriented post-training first.

For buyers, the practical question is not whether V4-Flash is the absolute best coding agent. It is whether a cheaper model with an 82.7 Terminal Bench 2.1 score can take the routine work that does not justify premium frontier inference. That is the part of the market most likely to move first.