GPT-56T 861 —
MUSE-SPK 837 +0.2%
GPT-56SC 790 -4.6%
GLM-5 781 -0.4%
CL-OP55X 780 -5.1%
GROK-46H 780 -5.1%
QWEN-38X 748 -9.2%
GPT-6A 743 -9.4%
KIMI-K3X 742 —
CL-FAB5H 698 -6.1%
CL-OP5H 675 -6.2%
GEM-38FH 672 -0.7%
CL-OP5X 670 -5.5%
CL-OP55H 668 —
CL-OP46H 657 -5.9%
CL-OP47H 648 -6.1%
GPT-56S 618 -0.6%
GEM-37FH 610 -7.2%
GEM-36FH 593 —
CL-OP48H 588 —
CL-OP47 581 -0.2%
GEM-35FH 580 —
GPT-55H 541 -7%
INKL 531 —
GEM-31P 512 -0.2%
CL-OP46 498 +0.4%
GEM-3P 498 -0.2%
CL-OP48 492 +0.4%
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 +0.2%
GPT-56SC 790 -4.6%
GLM-5 781 -0.4%
CL-OP55X 780 -5.1%
GROK-46H 780 -5.1%
QWEN-38X 748 -9.2%
GPT-6A 743 -9.4%
KIMI-K3X 742 —
CL-FAB5H 698 -6.1%
CL-OP5H 675 -6.2%
GEM-38FH 672 -0.7%
CL-OP5X 670 -5.5%
CL-OP55H 668 —
CL-OP46H 657 -5.9%
CL-OP47H 648 -6.1%
GPT-56S 618 -0.6%
GEM-37FH 610 -7.2%
GEM-36FH 593 —
CL-OP48H 588 —
CL-OP47 581 -0.2%
GEM-35FH 580 —
GPT-55H 541 -7%
INKL 531 —
GEM-31P 512 -0.2%
CL-OP46 498 +0.4%
GEM-3P 498 -0.2%
CL-OP48 492 +0.4%
GPT-52 464 —
GPT-55 423 —
← Back to feed

DeepSeek Puts Three V4 Variants on Arena Simultaneously — 80.6% SWE-Bench Puts Pro at the Frontier Tier

Chatbot Arena’s leaderboard changelog registered three DeepSeek entries on April 23: deepseek-v4-pro, deepseek-v4-pro-thinking, and deepseek-v4-flash-thinking. All three went live on both the Text and Code leaderboards simultaneously — an unusual batch entry that signals DeepSeek is aggressively pursuing full benchmark coverage for its V4 generation.

What’s Been Independently Verified

SWE-Bench Verified scores for the V4 family, confirmed from the tech report and independent evaluators:

ModelSWE-Bench Verified
DeepSeek V4 Pro80.6%
DeepSeek V4 Flash79.0%
DeepSeek V3 (prior gen)73.0%

V4 Pro’s 80.6% matches Gemini 3.1 Pro Preview’s 80.6% exactly, and trails Claude Opus 4.7’s 87.6% and Kimi K2.6’s 80.2% only slightly. The jump from V3 (73.0%) to V4 Pro (80.6%) is 7.6 points — comparable to the generational gains other labs have posted.

The Three-Variant Strategy

The simultaneous Arena entry with three variants suggests DeepSeek has structured V4 as a product family rather than a single release:

  • V4 Pro: Standard reasoning, 1.6T parameters, Apache 2.0
  • V4 Pro Thinking: Extended reasoning mode with chain-of-thought
  • V4 Flash Thinking: Lighter-weight with reasoning capability at V4 Flash costs ($0.14/M input)

This mirrors the pattern Anthropic uses with Haiku/Sonnet/Opus and their thinking variants — a cost-capability ladder built into the product line from day one.

The Cost Picture

V4 Flash launched at $0.14/M input, making it among the cheapest models at its capability tier. V4 Pro pricing has not been separately disclosed from V4 Flash in public rate sheets as of publication, but Artificial Analysis has a dedicated analysis page live for V4 Pro (Reasoning, Max Effort), suggesting pricing data is in their system.

Why This Batch Entry Matters

Labs with strong models typically submit to Arena over weeks as they fine-tune for benchmark conditions. Submitting three variants on the same day is aggressive positioning ahead of Arena ELO accumulation — it means all three will begin building vote counts from the same baseline date. The arena ELO positions for V4 Pro, V4 Pro Thinking, and V4 Flash Thinking will be visible once sufficient head-to-head votes accumulate, likely within 2–3 weeks.

The independent SWE data already confirms that V4 Pro belongs in the same tier as Gemini 3.1 Pro Preview and Kimi K2.6 — and 7 points ahead of where DeepSeek V3 sat.