GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
← Back to feed

DeepSeek V4-Flash Costs $0.03 per AA Task: 105x Cheaper Than Fable 5, Intelligence Index at 50

Artificial Analysis has completed its evaluation of DeepSeek V4-Flash 0731, and the headline number is cost: $0.03 per benchmark task, the lowest of any well-known model in the AA comparison set. Claude Fable 5 runs $3.15 per task, GPT-5.6 Sol $1.86, and Kimi K3 $0.86. On a per-task basis, V4-Flash costs 105 times less than Fable 5.

The cost-per-task metric accounts for the full amount of data a model processes and generates to complete a task — not just the headline token price. A model priced low but verbose can end up more expensive per completion. V4-Flash at $0.14/$0.28 per million tokens is already the cheapest headline price in the comparison, and it is also the most token-efficient on AA’s evaluation suite.

Intelligence Index Score: 50

That cost efficiency comes with a capability tradeoff. DeepSeek V4-Flash scores 50 on AA’s Intelligence Index v4.1, which combines nine benchmarks spanning coding, reasoning, and professional work tasks.

ModelAA Intelligence IndexAA Cost per Task
Claude Opus 5 / Fable 5 / GPT-5.659–66+$0.70–$3.15
Kimi K357$0.86
Meta Muse Spark 1.1 / GLM-5.251
DeepSeek V4-Flash 073150$0.03
Google Gemini 3.6 Flash50

The 7-point gap to Kimi K3 is the same spread that separates two consecutive Gemini Flash tiers. At the frontier, the gap to Fable 5 (64+ on current AA Index) is larger — roughly the distance from Fable 5 to a capable mid-tier model.

What 50 means in practice: V4-Flash operates at a level where most professional coding, reasoning, and analysis tasks complete correctly in straightforward cases. It degrades on complex multi-step problems and the hardest knowledge-intensive benchmarks. The AA evaluation shows it tied with Gemini 3.6 Flash, one point behind GLM-5.2 — both strong performers for their price tier.

Where DeepSeek Stands in the Chinese AI Field

DeepSeek launched V4-Flash in late July with 79% SWE-bench Verified on 13 billion active parameters — a significant efficiency result at its parameter count. Its AA score landing at 50 means it matches models with much larger active parameter counts on AA’s broad evaluation.

Among Chinese AI models with AA scores, Kimi K3 (57) holds a meaningful lead. GLM-5.2 (51) is one point ahead. Qwen3.8-Max does not yet have an AA score as of August 3. On Arena’s independent Frontend Code leaderboard, V4-Flash High came in with a preliminary score of 1,577 (1,319 votes), behind Qwen3.8-Max (1,668) and Kimi K3 (1,676) — consistent with the AA ranking order.

V4-Pro Next

DeepSeek has disclosed that V4-Pro is in preparation. No date has been given. The V4-Flash 0731 cycle — a smaller, cheaper, faster model — appears to be DeepSeek’s response to the Kimi K3 and GLM-5.2 launches that pushed past it on AA while DeepSeek held the cost position. The open question for V4-Pro is whether DeepSeek can close the intelligence gap to Kimi K3 without abandoning the cost advantage that defines its market position.

The Cost-to-Capability Ratio

At 105x cheaper than Fable 5 for 76–77% of Fable 5’s AA intelligence score (50 vs 66), V4-Flash offers the largest cost-to-capability ratio in the current evaluated field. The practical question is whether the 16-point gap to Fable 5 on AA affects production workloads in the way that price savings do. For high-volume, simple-to-moderate complexity tasks, the math strongly favors V4-Flash. For complex agentic work where model failures compound across many steps, the gap to Kimi K3 at $0.86/task — only 29x cheaper than Fable 5, but 7 AA points stronger — is the closer comparison.