GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
← Back to feed

LiveBench: Claude Fable 5 Max Effort Takes #1 at 83.0, Leads GPT-5.6 Sol by 2 Points

Claude Fable 5 Max Effort has taken first place on LiveBench with an overall score of 83.0, moving past GPT-5.6 Sol Max Effort at 81.0. GPT-5.6 Sol previously led at 82.4 in the last reported standings. The 2-point gap is narrow but consistent across multiple categories.

Current Top-5 Standings

ModelOverallAgentic CodingCost/Task
Claude Fable 5 Max Effort83.062.2$1.439
GPT-5.6 Sol Max Effort81.056.2$0.438
GPT-5.5 Thinking xHigh80.254.0$0.435
Claude Opus 5 Thinking Max80.165.2$0.699
Kimi K3 (open)79.262.2$0.348

Fable 5 leads in Reasoning (89.7), Coding (86.0), Language (90.7), and Mathematics (96.0). It ties with Kimi K3 on Agentic Coding at 62.2, where Claude Opus 5 Thinking Max Effort leads the table at 65.2.

The Cost-Efficiency Picture

Fable 5 Max Effort costs $1.439 per successful task, more than 3x what GPT-5.6 Sol charges ($0.438) for a result 2 points lower. Claude Opus 5 Thinking Max Effort at $0.699 sits 2.9 points below Fable 5 while costing less than half as much.

Kimi K3 makes the strongest cost argument: 79.2 overall at $0.348 per task. That’s 4.6% below Fable 5’s score at 24% of the cost. For organisations routing tasks to the cheapest model that can pass a quality threshold, K3 is the only open-weight option that lands inside the frontier tier.

What Agentic Coding Reveals

The Agentic Coding category, where models execute multi-step programming tasks across a real environment, shows a different ranking than the headline overall score. Claude Opus 5 Thinking Max Effort leads there at 65.2, with Fable 5 and Kimi K3 tied at 62.2, and GPT-5.6 Sol at 56.2.

GPT-5.6 Terra, previously reported leading Agentic Coding at 68.0% in an earlier LiveBench run, appears further down the current table at 77.9 overall. The overall composite includes categories where Terra is weaker, pulling its headline number below Sol.

Benchmark Context

LiveBench updates continuously as new model configurations are added, and results reflect the best published submission for each model and effort level, not necessarily default API settings. Max Effort configurations use extended thinking budgets and multiple passes. Users evaluating models at standard API settings should treat the per-task cost column as the operative variable.