GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

China Holds All 10 Spots on the Open-Weights Intelligence Index — Frontier Gap Collapsed From 13 to 6 Points in One Year

The top 10 open-weights models on Artificial Analysis’s Intelligence Index are all from China-based AI labs — not a near-miss, but a clean sweep. The analysis, published April 30, reflects results across Kimi K2.6 (Moonshot AI), MiMo V2.5 Pro (Xiaomi), DeepSeek V4 Pro (DeepSeek), and seven more models originating from Chinese labs. The only non-China models to crack the top tier: Google’s Gemma 4 31B and NVIDIA Nemotron 3 Super, sitting at positions 11 and 12.

The Numbers

ModelLabIntelligence Index
Kimi K2.6 (Reasoning)Moonshot AI54
MiMo V2.5 Pro (Reasoning)Xiaomi54
DeepSeek V4 Pro (Reasoning, Max)DeepSeek52
… (positions 4-10)China-based42-51
GPT-5.5 (xhigh)OpenAI60
Claude Opus 4.7 / Gemini 3.1 ProAnthropic / Google57

The gap from the best open-weights model to frontier proprietary: 6 Intelligence Index points.

One year ago, the best open-weights model was DeepSeek V3 0324 at score 22. The best proprietary model was Claude 3.7 Sonnet at 35. Gap: 13 points.

Architecture Profile

All three leading open-weights models are trillion-plus-parameter mixture-of-experts architectures with permissive licenses:

  • Kimi K2.6: 1T total / 32B active parameters, 256K context, Apache 2.0
  • MiMo V2.5 Pro: 1T total / 42B active, 1M context, Apache 2.0
  • DeepSeek V4 Pro: 1.6T total / 49B active, 1M context, MIT

The pattern is consistent: MoE architecture to spread parameters cheaply, sparse activation to keep inference costs manageable, permissive licensing to drive adoption and routing.

Where the Gap Remains Real

The Intelligence Index aggregate masks hard limits. On the most demanding evaluations, the gap to proprietary widens sharply:

HLE (Humanity’s Last Exam): Top open-weights score 34-36%. GPT-5.5 (xhigh): 44%. Gemini 3.1 Pro: 45%.

CritPt (Research-level physics): Open-weights 4-12%. GPT-5.5: 27%.

TerminalBench Hard (agentic coding): Open-weights 43-46%. GPT-5.5: 61%. Gemini 3.1 Pro: 54%.

Hallucination (Omniscience score): DeepSeek V4 Pro scores -10 (net negative on the hallucination-adjusted measure). MiMo V2.5 Pro: +4. Kimi K2.6: +6. GPT-5.5: +20. Claude Opus 4.7: +26. Gemini 3.1 Pro: +33.

The reliability gap is the structural problem. Intelligence Index performance has converged. Hallucination rates and hard-problem accuracy have not.

Price-Performance Picture

Open-weights models dominate the cost-efficiency frontier. Nine of the 13 models on Artificial Analysis’s Intelligence-vs-Price Pareto frontier are open-weights. Kimi K2.6 and MiMo V2.5 Pro both sit on the frontier; DeepSeek V4 Pro is just below. Across these three, users pay half to one-sixth the token cost of equivalent proprietary models for comparable Intelligence Index performance.

The Frontier Has Moved Since

The AA analysis reflects the state through late April 2026. Since then, Anthropic launched Claude Opus 4.8 on May 28, which leads the Intelligence Index at 61.4 with SWE-bench Verified at 88.6% — pushing the proprietary ceiling higher. The open-weights top scores have not materially shifted. The gap, already closed from 13 to 6 points in twelve months, has likely widened again by 1-2 points with Opus 4.8’s arrival.

The structural fact stands: the open-weights frontier is now a Chinese-origin product category, operating within striking distance of the best proprietary models on aggregate benchmarks, with a persistent but narrowing gap on the hardest tasks.