LiveBench: Claude Fable 5 Max Effort Takes #1 at 83.0, Leads GPT-5.6 Sol by 2 Points
Claude Fable 5 Max Effort has taken first place on LiveBench with an overall score of 83.0, moving past GPT-5.6 Sol Max Effort at 81.0. GPT-5.6 Sol previously led at 82.4 in the last reported standings. The 2-point gap is narrow but consistent across multiple categories.
Current Top-5 Standings
| Model | Overall | Agentic Coding | Cost/Task |
|---|---|---|---|
| Claude Fable 5 Max Effort | 83.0 | 62.2 | $1.439 |
| GPT-5.6 Sol Max Effort | 81.0 | 56.2 | $0.438 |
| GPT-5.5 Thinking xHigh | 80.2 | 54.0 | $0.435 |
| Claude Opus 5 Thinking Max | 80.1 | 65.2 | $0.699 |
| Kimi K3 (open) | 79.2 | 62.2 | $0.348 |
Fable 5 leads in Reasoning (89.7), Coding (86.0), Language (90.7), and Mathematics (96.0). It ties with Kimi K3 on Agentic Coding at 62.2, where Claude Opus 5 Thinking Max Effort leads the table at 65.2.
The Cost-Efficiency Picture
Fable 5 Max Effort costs $1.439 per successful task, more than 3x what GPT-5.6 Sol charges ($0.438) for a result 2 points lower. Claude Opus 5 Thinking Max Effort at $0.699 sits 2.9 points below Fable 5 while costing less than half as much.
Kimi K3 makes the strongest cost argument: 79.2 overall at $0.348 per task. That’s 4.6% below Fable 5’s score at 24% of the cost. For organisations routing tasks to the cheapest model that can pass a quality threshold, K3 is the only open-weight option that lands inside the frontier tier.
What Agentic Coding Reveals
The Agentic Coding category, where models execute multi-step programming tasks across a real environment, shows a different ranking than the headline overall score. Claude Opus 5 Thinking Max Effort leads there at 65.2, with Fable 5 and Kimi K3 tied at 62.2, and GPT-5.6 Sol at 56.2.
GPT-5.6 Terra, previously reported leading Agentic Coding at 68.0% in an earlier LiveBench run, appears further down the current table at 77.9 overall. The overall composite includes categories where Terra is weaker, pulling its headline number below Sol.
Benchmark Context
LiveBench updates continuously as new model configurations are added, and results reflect the best published submission for each model and effort level, not necessarily default API settings. Max Effort configurations use extended thinking budgets and multiple passes. Users evaluating models at standard API settings should treat the per-task cost column as the operative variable.