GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

LiveBench: GPT-5.6 Sol Leads at 82.4 as Kimi K3 Becomes Cheapest Frontier-Tier Model

The current LiveBench leaderboard has a new leader. GPT-5.6 Sol Max Effort posts 82.4 overall — above GPT-5.5 Thinking xHigh at 79.9, which held the top spot in the previous cycle. Fable 5 sits at 80.8, between them.

The bigger story is cost efficiency. At $0.379 per successful task, Kimi K3 matches Claude Opus 4.8 on overall score while running at less than 60% of GPT-5.5’s cost and less than 25% of Fable 5’s.

Top-5 Rankings

ModelOverallAgentic CodingCost/Task
GPT-5.6 Sol Max Effort82.465.6$0.589
Claude Fable 5 Max Effort80.846.9$1.573
GPT-5.5 Thinking xHigh79.952.1$0.530
GPT-5.6 Terra Max Effort79.868.0$0.497
Claude Opus 4.8 Thinking xHigh78.956.1$0.688
Kimi K3 (open-weight)78.557.6$0.379

Two Variants, Different Jobs

GPT-5.6 Sol and Terra diverge in exactly one meaningful dimension. Sol leads overall at 82.4; Terra leads Agentic Coding at 68.0. Terra also costs less per task ($0.497 vs $0.589). For teams running coding agents at scale, Terra is the rational choice — higher on the category that matters most, cheaper per run.

The Sol/Terra split mirrors a pattern seen with GPT-5.5 High vs xHigh: different effort tiers produce different capability profiles, not just scaled versions of the same profile. The distinction matters when selecting a model for a specific task class rather than general use.

Fable 5’s Agentic Coding Problem

Claude Fable 5’s overall score of 80.8 masks a weak result in the agentic coding subcategory: 46.9. That’s below GPT-5.5, Kimi K3, and Claude Opus 4.8. For a model that holds the #1 spot on the Artificial Analysis Intelligence Index and SWE-bench Verified, a 46.9 on LiveBench Agentic Coding is a notable gap — and the most expensive option on the leaderboard at $1.573/task.

Fable 5’s strengths are in language (90.7) and math (96.0). The benchmark split matters: use Fable 5 for reasoning-intensive tasks, not for long agentic coding runs.

Kimi K3’s Position

At 78.5 overall and $0.379 per task, Kimi K3 is the practical choice at the frontier tier for cost-constrained deployments. It leads Kimi K3’s open-weight peers by a substantial margin and beats every closed proprietary model below it on the leaderboard on the only dimension that’s ultimately measurable at production scale: cost per correct output.

The model is available as a 2.8T-parameter open-weight release, meaning the benchmark result is reproducible and self-hostable. Kimi K3 also leads Reasoning at 90.7 — tied with GPT-5.5 Thinking xHigh, above Sol.

What Changed From Last Cycle

The previous LiveBench coverage showed GPT-5.5 at the top and Claude Opus 4.8 leading Agentic Coding at 56.1. The current snapshot adds two GPT-5.6 variants above both, and inserts Kimi K3 into the frontier cluster. GPT-5.6 Terra’s 68.0 on Agentic Coding is the largest single-category gain relative to the previous leaderboard snapshot — a 12-point jump over Opus 4.8’s 56.1.