GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

DeepSeek V4-Pro 0813 Hits 96.4% on Neutral SWE-Bench — Open Weights at #2, 59x Cheaper Than Claude Opus 5

When DeepSeek published its V4 Pro tech report in April, it posted an 80.6% SWE-bench Verified score. That number has been the reference point for every competitive analysis since. It is wrong, in the sense that it reflects DeepSeek’s own test harness rather than a standardised one. On vals.ai’s neutral bash-only setup, the same model scores 96.40% — a 15.8-point gap that lands V4-Pro-0813 at second place globally.

The current top four on vals.ai’s SWE-bench Verified leaderboard, as of August 13:

RankModelScoreCost/task
1Claude Opus 597.00%$1.29
2DeepSeek V4-Pro-081396.40% ±0.83$0.022
3GPT-5.6 Sol96.20%
4Grok 4.695.60%$0.785
5DeepSeek V4-Flash-073188.80%$0.010

The gap between first and fourth is 1.4 percentage points. The gap between the top open-weight model and the top closed-weight model is 0.6 points.

Why the Numbers Differ

SWE-bench Verified scores are harness-dependent. The benchmark specifies tasks and scoring, not the scaffolding an agent uses to attempt them. DeepSeek’s April report used its own harness in high-creativity mode. Vals.ai runs every model through mini-SWE-agent in a minimal bash environment with no custom scaffolding — the same conditions for every model on the table.

The 15.8-point delta is not evidence that the model improved. It is evidence that harness choice matters as much as the weights when comparing published numbers across labs.

The model weights for V4-Pro-0813 are unchanged from the April preview. DeepSeek confirmed on July 31 that the Pro API would be promoted to GA without a weights update, which it was on August 12.

The Cost Picture

At $0.435 per million input tokens and $0.87 per million output, V4-Pro-0813 costs approximately $0.022 per SWE-bench task on a neutral harness. Claude Opus 5 runs $1.29 per task — 59 times higher. GPT-5.6 Sol does not have a published per-task figure, but its $30 per million output price puts blended costs well above $1.00 per task at typical benchmark workloads.

Across nine agent benchmarks where DeepSeek’s own table compares V4-Pro against Claude Fable 5 (a closer frontier peer than Opus 5), the average gap is 5.3% in Fable 5’s favour. The cost gap is 4,600%.

Grok 4.6 at 95.6%

Grok 4.6, released August 12, also appears in the top four on the neutral harness at 95.60%. SpaceXAI’s launch post highlighted its Artificial Analysis Intelligence Index score (61, tied with GPT-5.6 Sol Max) and CursorBench (69.9%), neither of which is directly comparable to SWE-bench. The 95.60% SWE figure was published independently by vals.ai.

This puts three models within 1.4 points of each other on the most widely cited coding benchmark, and an open-weight model inside that cluster at a price that most frontier competitors cannot match without a complete cost architecture redesign.