DeepSeek V4-Flash Costs $0.03 per AA Task: 105x Cheaper Than Fable 5, Intelligence Index at 50
Artificial Analysis has completed its evaluation of DeepSeek V4-Flash 0731, and the headline number is cost: $0.03 per benchmark task, the lowest of any well-known model in the AA comparison set. Claude Fable 5 runs $3.15 per task, GPT-5.6 Sol $1.86, and Kimi K3 $0.86. On a per-task basis, V4-Flash costs 105 times less than Fable 5.
The cost-per-task metric accounts for the full amount of data a model processes and generates to complete a task — not just the headline token price. A model priced low but verbose can end up more expensive per completion. V4-Flash at $0.14/$0.28 per million tokens is already the cheapest headline price in the comparison, and it is also the most token-efficient on AA’s evaluation suite.
Intelligence Index Score: 50
That cost efficiency comes with a capability tradeoff. DeepSeek V4-Flash scores 50 on AA’s Intelligence Index v4.1, which combines nine benchmarks spanning coding, reasoning, and professional work tasks.
| Model | AA Intelligence Index | AA Cost per Task |
|---|---|---|
| Claude Opus 5 / Fable 5 / GPT-5.6 | 59–66+ | $0.70–$3.15 |
| Kimi K3 | 57 | $0.86 |
| Meta Muse Spark 1.1 / GLM-5.2 | 51 | — |
| DeepSeek V4-Flash 0731 | 50 | $0.03 |
| Google Gemini 3.6 Flash | 50 | — |
The 7-point gap to Kimi K3 is the same spread that separates two consecutive Gemini Flash tiers. At the frontier, the gap to Fable 5 (64+ on current AA Index) is larger — roughly the distance from Fable 5 to a capable mid-tier model.
What 50 means in practice: V4-Flash operates at a level where most professional coding, reasoning, and analysis tasks complete correctly in straightforward cases. It degrades on complex multi-step problems and the hardest knowledge-intensive benchmarks. The AA evaluation shows it tied with Gemini 3.6 Flash, one point behind GLM-5.2 — both strong performers for their price tier.
Where DeepSeek Stands in the Chinese AI Field
DeepSeek launched V4-Flash in late July with 79% SWE-bench Verified on 13 billion active parameters — a significant efficiency result at its parameter count. Its AA score landing at 50 means it matches models with much larger active parameter counts on AA’s broad evaluation.
Among Chinese AI models with AA scores, Kimi K3 (57) holds a meaningful lead. GLM-5.2 (51) is one point ahead. Qwen3.8-Max does not yet have an AA score as of August 3. On Arena’s independent Frontend Code leaderboard, V4-Flash High came in with a preliminary score of 1,577 (1,319 votes), behind Qwen3.8-Max (1,668) and Kimi K3 (1,676) — consistent with the AA ranking order.
V4-Pro Next
DeepSeek has disclosed that V4-Pro is in preparation. No date has been given. The V4-Flash 0731 cycle — a smaller, cheaper, faster model — appears to be DeepSeek’s response to the Kimi K3 and GLM-5.2 launches that pushed past it on AA while DeepSeek held the cost position. The open question for V4-Pro is whether DeepSeek can close the intelligence gap to Kimi K3 without abandoning the cost advantage that defines its market position.
The Cost-to-Capability Ratio
At 105x cheaper than Fable 5 for 76–77% of Fable 5’s AA intelligence score (50 vs 66), V4-Flash offers the largest cost-to-capability ratio in the current evaluated field. The practical question is whether the 16-point gap to Fable 5 on AA affects production workloads in the way that price savings do. For high-volume, simple-to-moderate complexity tasks, the math strongly favors V4-Flash. For complex agentic work where model failures compound across many steps, the gap to Kimi K3 at $0.86/task — only 29x cheaper than Fable 5, but 7 AA points stronger — is the closer comparison.