Kimi K2.6 Hits Intelligence Index 54, Three Points Behind the Frontier Ceiling
Moonshot AI’s Kimi K2.6 has landed at #4 on the Artificial Analysis Intelligence Index with a score of 54, three points behind the three-way tie at 57 held by Claude Opus 4.7, Gemini 3.1 Pro Preview, and GPT-5.4. It is the only open-weights model in the top five.
The result marks a meaningful jump from Kimi K2.5. On GDPval-AA — AA’s primary agentic capability metric, which tests knowledge-work tasks including presentations and analysis with code execution and web browsing tools in a live loop — K2.6 posts an Elo of 1,520, compared to K2.5’s 1,309. That 211-point gain is the sharpest generation-over-generation agentic improvement in the current top tier.
Key Numbers
- AA Intelligence Index: 54 (vs K2.5’s unranked)
- GDPval-AA Elo: 1,520 (K2.5: 1,309, +211)
- τ²-Bench Telecom: 96% — in line with Claude Opus 4.7 and GPT-5.4
- Hallucination rate: 39% (down from 65% in K2.5)
- Parameters: 1T total, 32B active (Mixture-of-Experts)
- Context window: 256K tokens
- Reasoning token usage: ~160M for full AA eval (vs Claude Sonnet 4.6’s ~190M, GPT-5.4’s ~110M)
The Three-Point Gap
The 57-vs-54 split between proprietary frontier and open weights is the tightest it has been. For context, Kimi K2.5 sat well outside the top tier; the two Gemma 4 models with the best AA scores land in the 30s. K2.6 closing to within three points — on a public, Apache-licensed model — is the current high-water mark for open-weights capability.
The gap is not evenly distributed across tasks. K2.6’s hallucination rate of 39% places it comparably to Claude Opus 4.7 (36%) and MiniMax-M2.7 (34%), which is notable: most open-weights models hallucinate at significantly higher rates than their intelligence scores would suggest.
Where K2.6 still trails: reasoning token efficiency. GPT-5.4 runs the full AA Index on roughly 110M reasoning tokens; K2.6 uses 160M. The model is more expensive to run at inference than its proprietary peers at the same output quality, which partially offsets the cost advantages of open weights.
Access
K2.6 is available through Moonshot’s first-party API and via Novita, Baseten, Fireworks, and Parasail. The model supports image and video input with text output, with a 256K max context.