GPT-56T 861 —
MUSE-SPK 837 +0.2%
GPT-56SC 790 -4.6%
GLM-5 781 -0.4%
CL-OP55X 780 -5.1%
GROK-46H 780 -5.1%
QWEN-38X 748 -9.2%
GPT-6A 743 -9.4%
KIMI-K3X 742 —
CL-FAB5H 698 -6.1%
CL-OP5H 675 -6.2%
GEM-38FH 672 -0.7%
CL-OP5X 670 -5.5%
CL-OP55H 668 —
CL-OP46H 657 -5.9%
CL-OP47H 648 -6.1%
GPT-56S 618 -0.6%
GEM-37FH 610 -7.2%
GEM-36FH 593 —
CL-OP48H 588 —
CL-OP47 581 -0.2%
GEM-35FH 580 —
GPT-55H 541 -7%
INKL 531 —
GEM-31P 512 -0.2%
CL-OP46 498 +0.4%
GEM-3P 498 -0.2%
CL-OP48 492 +0.4%
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 +0.2%
GPT-56SC 790 -4.6%
GLM-5 781 -0.4%
CL-OP55X 780 -5.1%
GROK-46H 780 -5.1%
QWEN-38X 748 -9.2%
GPT-6A 743 -9.4%
KIMI-K3X 742 —
CL-FAB5H 698 -6.1%
CL-OP5H 675 -6.2%
GEM-38FH 672 -0.7%
CL-OP5X 670 -5.5%
CL-OP55H 668 —
CL-OP46H 657 -5.9%
CL-OP47H 648 -6.1%
GPT-56S 618 -0.6%
GEM-37FH 610 -7.2%
GEM-36FH 593 —
CL-OP48H 588 —
CL-OP47 581 -0.2%
GEM-35FH 580 —
GPT-55H 541 -7%
INKL 531 —
GEM-31P 512 -0.2%
CL-OP46 498 +0.4%
GEM-3P 498 -0.2%
CL-OP48 492 +0.4%
GPT-52 464 —
GPT-55 423 —
← Back to feed

Kimi K2.6 Hits Intelligence Index 54, Three Points Behind the Frontier Ceiling

Moonshot AI’s Kimi K2.6 has landed at #4 on the Artificial Analysis Intelligence Index with a score of 54, three points behind the three-way tie at 57 held by Claude Opus 4.7, Gemini 3.1 Pro Preview, and GPT-5.4. It is the only open-weights model in the top five.

The result marks a meaningful jump from Kimi K2.5. On GDPval-AA — AA’s primary agentic capability metric, which tests knowledge-work tasks including presentations and analysis with code execution and web browsing tools in a live loop — K2.6 posts an Elo of 1,520, compared to K2.5’s 1,309. That 211-point gain is the sharpest generation-over-generation agentic improvement in the current top tier.

Key Numbers

  • AA Intelligence Index: 54 (vs K2.5’s unranked)
  • GDPval-AA Elo: 1,520 (K2.5: 1,309, +211)
  • τ²-Bench Telecom: 96% — in line with Claude Opus 4.7 and GPT-5.4
  • Hallucination rate: 39% (down from 65% in K2.5)
  • Parameters: 1T total, 32B active (Mixture-of-Experts)
  • Context window: 256K tokens
  • Reasoning token usage: ~160M for full AA eval (vs Claude Sonnet 4.6’s ~190M, GPT-5.4’s ~110M)

The Three-Point Gap

The 57-vs-54 split between proprietary frontier and open weights is the tightest it has been. For context, Kimi K2.5 sat well outside the top tier; the two Gemma 4 models with the best AA scores land in the 30s. K2.6 closing to within three points — on a public, Apache-licensed model — is the current high-water mark for open-weights capability.

The gap is not evenly distributed across tasks. K2.6’s hallucination rate of 39% places it comparably to Claude Opus 4.7 (36%) and MiniMax-M2.7 (34%), which is notable: most open-weights models hallucinate at significantly higher rates than their intelligence scores would suggest.

Where K2.6 still trails: reasoning token efficiency. GPT-5.4 runs the full AA Index on roughly 110M reasoning tokens; K2.6 uses 160M. The model is more expensive to run at inference than its proprietary peers at the same output quality, which partially offsets the cost advantages of open weights.

Access

K2.6 is available through Moonshot’s first-party API and via Novita, Baseten, Fireworks, and Parasail. The model supports image and video input with text output, with a 256K max context.