DeepSeek Puts Three V4 Variants on Arena Simultaneously — 80.6% SWE-Bench Puts Pro at the Frontier Tier
Chatbot Arena’s leaderboard changelog registered three DeepSeek entries on April 23: deepseek-v4-pro, deepseek-v4-pro-thinking, and deepseek-v4-flash-thinking. All three went live on both the Text and Code leaderboards simultaneously — an unusual batch entry that signals DeepSeek is aggressively pursuing full benchmark coverage for its V4 generation.
What’s Been Independently Verified
SWE-Bench Verified scores for the V4 family, confirmed from the tech report and independent evaluators:
| Model | SWE-Bench Verified |
|---|---|
| DeepSeek V4 Pro | 80.6% |
| DeepSeek V4 Flash | 79.0% |
| DeepSeek V3 (prior gen) | 73.0% |
V4 Pro’s 80.6% matches Gemini 3.1 Pro Preview’s 80.6% exactly, and trails Claude Opus 4.7’s 87.6% and Kimi K2.6’s 80.2% only slightly. The jump from V3 (73.0%) to V4 Pro (80.6%) is 7.6 points — comparable to the generational gains other labs have posted.
The Three-Variant Strategy
The simultaneous Arena entry with three variants suggests DeepSeek has structured V4 as a product family rather than a single release:
- V4 Pro: Standard reasoning, 1.6T parameters, Apache 2.0
- V4 Pro Thinking: Extended reasoning mode with chain-of-thought
- V4 Flash Thinking: Lighter-weight with reasoning capability at V4 Flash costs ($0.14/M input)
This mirrors the pattern Anthropic uses with Haiku/Sonnet/Opus and their thinking variants — a cost-capability ladder built into the product line from day one.
The Cost Picture
V4 Flash launched at $0.14/M input, making it among the cheapest models at its capability tier. V4 Pro pricing has not been separately disclosed from V4 Flash in public rate sheets as of publication, but Artificial Analysis has a dedicated analysis page live for V4 Pro (Reasoning, Max Effort), suggesting pricing data is in their system.
Why This Batch Entry Matters
Labs with strong models typically submit to Arena over weeks as they fine-tune for benchmark conditions. Submitting three variants on the same day is aggressive positioning ahead of Arena ELO accumulation — it means all three will begin building vote counts from the same baseline date. The arena ELO positions for V4 Pro, V4 Pro Thinking, and V4 Flash Thinking will be visible once sufficient head-to-head votes accumulate, likely within 2–3 weeks.
The independent SWE data already confirms that V4 Pro belongs in the same tier as Gemini 3.1 Pro Preview and Kimi K2.6 — and 7 points ahead of where DeepSeek V3 sat.