GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Kimi K3's First Benchmark Numbers: 90.7 LiveBench Reasoning, #4 on AA Intelligence Index at $0.38 Per Task

Kimi K3 launched July 16 without benchmark data. The numbers are now in. Moonshot AI’s 2.8-trillion-parameter open-weight model places 6th overall on LiveBench at 78.5 — and lands in the top two globally on reasoning at 90.7, one point behind GPT-5.6 Sol Max Effort (91.7) and above Claude Fable 5 (89.7).

LiveBench Scores

ModelOverallReasoningCodingAgentic CodingCost/Task
GPT-5.6 Sol Max Effort82.491.783.965.6$0.589
Claude Fable 5 Max Effort80.889.786.046.9$1.573
GPT-5.5 Thinking xHigh79.989.782.152.1$0.530
GPT-5.6 Terra Max Effort79.890.678.268.0$0.497
Claude Opus 4.8 xHigh78.989.779.356.1$0.688
Kimi K3 open78.590.781.457.6$0.379

Three findings stand out.

Reasoning is the outlier. At 90.7, K3 is the second-highest reasoning score on LiveBench — a 1.0-point gap behind GPT-5.6 Sol and matching GPT-5.6 Terra, both proprietary models from OpenAI. Claude Fable 5, which leads on coding (86.0) and language tasks, scores lower on reasoning at 89.7.

Cost per task is structurally different. K3 at $0.379 per successful task is 36% cheaper than the next-cheapest model in the top six (GPT-5.6 Terra at $0.497) and 4.15x cheaper than Fable 5 ($1.573). In workloads where throughput matters more than maximum accuracy, the economics of using K3 over Fable 5 are difficult to argue against.

Agentic coding is the weak point. K3 scores 57.6 on LiveBench agentic coding — behind GPT-5.6 Terra (68.0), GPT-5.6 Sol (65.6), and Opus 4.8 (56.1, with K3 slightly above). This puts it behind the coding-specialist tier despite leading on pure reasoning. The launch article for Kimi K3 flagged that its architecture (KDA + Attention Residuals) was designed for long agentic runs; the data suggests the optimization hasn’t fully closed the gap with GPT-5.6 Terra on agentic task completion.

Artificial Analysis Intelligence Index

On Artificial Analysis’s Intelligence Index — the broadest quality composite across frontier models — K3 ranks 4th globally, behind Claude Fable 5 (with fallback), GPT-5.6 Sol (max), and GPT-5.6 Sol (xhigh). That positions an open-weight model, accessible via downloadable weights on HuggingFace, inside the acknowledged top 4 on the world’s most comprehensive model leaderboard.

The prior K-series ceiling was Kimi K2.6 at 44 on the AA index. Kimi K3’s #4 placement across all model categories is a material step up.

What the Numbers Mean

Kimi K3’s 90.7 reasoning result disrupts a narrative that had hardened since Fable 5’s launch: that the frontier reasoning gap between open-weight and proprietary models was fixed. It is not. On reasoning — the benchmark dimension that most directly predicts performance on multi-step problem-solving — K3 is effectively at parity with every non-Sol proprietary model.

The agentic coding gap is real and worth watching. GPT-5.6 Terra leads that category at 68.0; K3 at 57.6 is 10.4 points back. If Moonshot AI’s post-launch tuning closes that specific deficit, K3 at $3/M input and $15/M output becomes a structural challenge to the proprietary agentic tier.

At $0.38 per successful task, K3 sits in the same cohort as GPT-5.6 models on overall LiveBench performance. The models above it all cost more. That is the open-weight argument in precise, current numbers.