Kimi K3's First Benchmark Numbers: 90.7 LiveBench Reasoning, #4 on AA Intelligence Index at $0.38 Per Task
Kimi K3 launched July 16 without benchmark data. The numbers are now in. Moonshot AI’s 2.8-trillion-parameter open-weight model places 6th overall on LiveBench at 78.5 — and lands in the top two globally on reasoning at 90.7, one point behind GPT-5.6 Sol Max Effort (91.7) and above Claude Fable 5 (89.7).
LiveBench Scores
| Model | Overall | Reasoning | Coding | Agentic Coding | Cost/Task |
|---|---|---|---|---|---|
| GPT-5.6 Sol Max Effort | 82.4 | 91.7 | 83.9 | 65.6 | $0.589 |
| Claude Fable 5 Max Effort | 80.8 | 89.7 | 86.0 | 46.9 | $1.573 |
| GPT-5.5 Thinking xHigh | 79.9 | 89.7 | 82.1 | 52.1 | $0.530 |
| GPT-5.6 Terra Max Effort | 79.8 | 90.6 | 78.2 | 68.0 | $0.497 |
| Claude Opus 4.8 xHigh | 78.9 | 89.7 | 79.3 | 56.1 | $0.688 |
| Kimi K3 open | 78.5 | 90.7 | 81.4 | 57.6 | $0.379 |
Three findings stand out.
Reasoning is the outlier. At 90.7, K3 is the second-highest reasoning score on LiveBench — a 1.0-point gap behind GPT-5.6 Sol and matching GPT-5.6 Terra, both proprietary models from OpenAI. Claude Fable 5, which leads on coding (86.0) and language tasks, scores lower on reasoning at 89.7.
Cost per task is structurally different. K3 at $0.379 per successful task is 36% cheaper than the next-cheapest model in the top six (GPT-5.6 Terra at $0.497) and 4.15x cheaper than Fable 5 ($1.573). In workloads where throughput matters more than maximum accuracy, the economics of using K3 over Fable 5 are difficult to argue against.
Agentic coding is the weak point. K3 scores 57.6 on LiveBench agentic coding — behind GPT-5.6 Terra (68.0), GPT-5.6 Sol (65.6), and Opus 4.8 (56.1, with K3 slightly above). This puts it behind the coding-specialist tier despite leading on pure reasoning. The launch article for Kimi K3 flagged that its architecture (KDA + Attention Residuals) was designed for long agentic runs; the data suggests the optimization hasn’t fully closed the gap with GPT-5.6 Terra on agentic task completion.
Artificial Analysis Intelligence Index
On Artificial Analysis’s Intelligence Index — the broadest quality composite across frontier models — K3 ranks 4th globally, behind Claude Fable 5 (with fallback), GPT-5.6 Sol (max), and GPT-5.6 Sol (xhigh). That positions an open-weight model, accessible via downloadable weights on HuggingFace, inside the acknowledged top 4 on the world’s most comprehensive model leaderboard.
The prior K-series ceiling was Kimi K2.6 at 44 on the AA index. Kimi K3’s #4 placement across all model categories is a material step up.
What the Numbers Mean
Kimi K3’s 90.7 reasoning result disrupts a narrative that had hardened since Fable 5’s launch: that the frontier reasoning gap between open-weight and proprietary models was fixed. It is not. On reasoning — the benchmark dimension that most directly predicts performance on multi-step problem-solving — K3 is effectively at parity with every non-Sol proprietary model.
The agentic coding gap is real and worth watching. GPT-5.6 Terra leads that category at 68.0; K3 at 57.6 is 10.4 points back. If Moonshot AI’s post-launch tuning closes that specific deficit, K3 at $3/M input and $15/M output becomes a structural challenge to the proprietary agentic tier.
At $0.38 per successful task, K3 sits in the same cohort as GPT-5.6 models on overall LiveBench performance. The models above it all cost more. That is the open-weight argument in precise, current numbers.