GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
← Back to feed

Kimi K3 Ties Fable 5 on LiveBench Agentic Coding at 62.2% and Costs $0.35 vs $1.44 Per Task

The current LiveBench leaderboard shows an uncomfortable tie at the frontier: Moonshot AI’s Kimi K3 and Anthropic’s Fable 5 both score 62.2% on agentic coding. They score identically. Kimi K3 costs $0.348 per successful task. Fable 5 costs $1.439.

That is a 4.1x cost difference for zero performance difference on the benchmark that most directly measures production agent capability.

Full Leaderboard State

ModelOverallAgentic CodingCost/Task
Claude Fable 5 Max Effort83.062.2%$1.439
GPT-5.6 Sol Max Effort81.056.2%$0.438
GPT-5.5 Thinking xHigh80.254.0%$0.435
Claude Opus 5 Thinking Max80.165.2%$0.699
Kimi K379.262.2%$0.348
GPT-5.4 Thinking xHigh78.053.8%$0.387
GPT-5.6 Terra Max Effort77.968.0%*

*Terra’s agentic coding score from previous benchmark run; cost not listed in current table.

What the Numbers Say

Fable 5 leads the overall LiveBench at 83.0 — nearly 4 points ahead of the next model — because it dominates reasoning (89.7%), coding (86.0%), mathematics (96.0%), and language (90.7%). On those dimensions, it is the strongest model available. But agentic coding is a different task. It rewards persistence, tool use under ambiguity, and multi-step problem decomposition — capabilities that are increasingly learnable through post-training, not just scale.

Kimi K3’s 62.2% on agentic coding comes from a 2.8 trillion parameter open-weight model running at $3/$15 per million tokens. It already cracked Agent Arena in July at 1543 AA-Briefcase Elo. Its agentic coding performance matching Fable 5 is not a fluke of one benchmark — it is consistent across evaluations.

Claude Opus 5 leads both at 65.2% and at $0.699 per task, which is the price-performance frontier on this specific dimension: 3 points better than Fable 5 and Kimi K3, at roughly half Fable 5’s task cost and double Kimi K3’s.

The Commoditization Arc

A year ago, frontier agentic coding capability was the exclusive property of a handful of closed models. Now an open-weight 2.8T MoE model matches Anthropic’s flagship on the benchmark that matters most for enterprise coding agents — and does it at a price that makes Fable 5 look like a premium tier with no performance premium to show for it.

For developers choosing a model for agentic coding pipelines: Kimi K3 at $0.35 per successful task, Opus 5 at $0.70 for a 3-point edge, or Fable 5 at $1.44 for the same score as Kimi K3. The overall LiveBench lead Fable 5 holds (83.0 vs 79.2) is real and matters for multi-domain workloads. For pure agentic coding, the spreadsheet does not favor it.

GPT-5.6 Terra leads agentic coding at 68.0% but at pricing that the current table does not list — and it sits 5 points below the Opus 5 overall score, with a weaker showing on reasoning and math. For most enterprise agent deployments, Opus 5 or Kimi K3 represent the current efficiency frontier.

Fable 5 is still the best overall model. It is no longer the best agentic coding value.