China Holds All 10 Spots on the Open-Weights Intelligence Index — Frontier Gap Collapsed From 13 to 6 Points in One Year
The top 10 open-weights models on Artificial Analysis’s Intelligence Index are all from China-based AI labs — not a near-miss, but a clean sweep. The analysis, published April 30, reflects results across Kimi K2.6 (Moonshot AI), MiMo V2.5 Pro (Xiaomi), DeepSeek V4 Pro (DeepSeek), and seven more models originating from Chinese labs. The only non-China models to crack the top tier: Google’s Gemma 4 31B and NVIDIA Nemotron 3 Super, sitting at positions 11 and 12.
The Numbers
| Model | Lab | Intelligence Index |
|---|---|---|
| Kimi K2.6 (Reasoning) | Moonshot AI | 54 |
| MiMo V2.5 Pro (Reasoning) | Xiaomi | 54 |
| DeepSeek V4 Pro (Reasoning, Max) | DeepSeek | 52 |
| … (positions 4-10) | China-based | 42-51 |
| GPT-5.5 (xhigh) | OpenAI | 60 |
| Claude Opus 4.7 / Gemini 3.1 Pro | Anthropic / Google | 57 |
The gap from the best open-weights model to frontier proprietary: 6 Intelligence Index points.
One year ago, the best open-weights model was DeepSeek V3 0324 at score 22. The best proprietary model was Claude 3.7 Sonnet at 35. Gap: 13 points.
Architecture Profile
All three leading open-weights models are trillion-plus-parameter mixture-of-experts architectures with permissive licenses:
- Kimi K2.6: 1T total / 32B active parameters, 256K context, Apache 2.0
- MiMo V2.5 Pro: 1T total / 42B active, 1M context, Apache 2.0
- DeepSeek V4 Pro: 1.6T total / 49B active, 1M context, MIT
The pattern is consistent: MoE architecture to spread parameters cheaply, sparse activation to keep inference costs manageable, permissive licensing to drive adoption and routing.
Where the Gap Remains Real
The Intelligence Index aggregate masks hard limits. On the most demanding evaluations, the gap to proprietary widens sharply:
HLE (Humanity’s Last Exam): Top open-weights score 34-36%. GPT-5.5 (xhigh): 44%. Gemini 3.1 Pro: 45%.
CritPt (Research-level physics): Open-weights 4-12%. GPT-5.5: 27%.
TerminalBench Hard (agentic coding): Open-weights 43-46%. GPT-5.5: 61%. Gemini 3.1 Pro: 54%.
Hallucination (Omniscience score): DeepSeek V4 Pro scores -10 (net negative on the hallucination-adjusted measure). MiMo V2.5 Pro: +4. Kimi K2.6: +6. GPT-5.5: +20. Claude Opus 4.7: +26. Gemini 3.1 Pro: +33.
The reliability gap is the structural problem. Intelligence Index performance has converged. Hallucination rates and hard-problem accuracy have not.
Price-Performance Picture
Open-weights models dominate the cost-efficiency frontier. Nine of the 13 models on Artificial Analysis’s Intelligence-vs-Price Pareto frontier are open-weights. Kimi K2.6 and MiMo V2.5 Pro both sit on the frontier; DeepSeek V4 Pro is just below. Across these three, users pay half to one-sixth the token cost of equivalent proprietary models for comparable Intelligence Index performance.
The Frontier Has Moved Since
The AA analysis reflects the state through late April 2026. Since then, Anthropic launched Claude Opus 4.8 on May 28, which leads the Intelligence Index at 61.4 with SWE-bench Verified at 88.6% — pushing the proprietary ceiling higher. The open-weights top scores have not materially shifted. The gap, already closed from 13 to 6 points in twelve months, has likely widened again by 1-2 points with Opus 4.8’s arrival.
The structural fact stands: the open-weights frontier is now a Chinese-origin product category, operating within striking distance of the best proprietary models on aggregate benchmarks, with a persistent but narrowing gap on the hardest tasks.