Claude Opus 5 Tops LiveBench Agentic Coding at 65.2% — Fable 5 Trails by 3 Points at Twice the Cost
The dominant pattern in frontier AI benchmarking has been Fable 5 holding the top line on most overall scores. LiveBench’s agentic coding dimension breaks that pattern. Claude Opus 5 Thinking Max Effort scores 65.2% on agentic coding, three points above Fable 5 Max Effort’s 62.2% and nine points above GPT-5.6 Sol’s 56.2%.
Across the full LiveBench leaderboard, Fable 5 leads at 83.0 overall. Opus 5 sits fourth at 80.1. But the models sort differently when the task is specifically agentic coding: Opus 5 moves to first, Kimi K3 ties Fable 5 at 62.2% in second, and Sol falls to fifth.
The Cost Math
The score difference is notable. The cost difference is sharper:
| Model | Agentic Coding | Cost Per Task |
|---|---|---|
| Claude Opus 5 Max Effort | 65.2% | $0.699 |
| Claude Fable 5 Max Effort | 62.2% | $1.439 |
| Kimi K3 | 62.2% | $0.348 |
| GPT-5.5 xHigh | 54.0% | $0.435 |
| GPT-5.6 Sol Max Effort | 56.2% | $0.438 |
Opus 5 beats Fable 5 on the metric and costs exactly half as much per successful task. Kimi K3 is the cheapest frontier-class option and ties Fable 5 at 62.2%; it loses to Opus 5 on score but undercuts it on cost at $0.348.
Why Agentic Coding Separates Here
LiveBench’s overall scores are dominated by reasoning, mathematics, and data analysis. These dimensions reward raw intelligence, where Fable 5’s training leads. Agentic coding requires a different profile: multi-step execution, tool use, context retention across long code sessions, and self-correction when initial approaches fail.
Opus 5 was positioned from its launch as a research and complex-task model. Its ARC-AGI-3 score of 30.2% — four times the previous best — reflects the same strength: sustained multi-step problem-solving that degrades slowly as session length increases. The LiveBench agentic coding result is independent confirmation of that design decision. On the dimension most closely resembling a real software engineering session, Opus 5 is the performance leader.
Artificial Analysis Cross-Check
The result aligns with Artificial Analysis data published this week. On the AA Intelligence Index, Opus 5 Adaptive Reasoning Max Effort scores 61, two points above GPT-5.6 Sol at 59. AA pricing data puts Opus 5 at $3.85 per million tokens blended versus Sol at $4.35, a 12% cost advantage. Time-to-first-token: Opus 5 at 65 seconds versus Sol at 143 seconds, a 2.2x response-time lead.
The two sources cross-reference cleanly. On AA Intelligence Index, Opus 5 beats Sol. On LiveBench agentic coding, Opus 5 beats Fable 5. On cost per successful task, Opus 5 undercuts Fable 5 by 51%.
Full LiveBench Top-5 Overview
| Model | Overall | Reasoning | Agentic Coding | Cost/Task |
|---|---|---|---|---|
| Fable 5 Max Effort | 83.0 | 89.7 | 62.2% | $1.439 |
| GPT-5.6 Sol Max Effort | 81.0 | 91.7 | 56.2% | $0.438 |
| GPT-5.5 xHigh | 80.2 | 89.7 | 54.0% | $0.435 |
| Claude Opus 5 Max Effort | 80.1 | 91.2 | 65.2% | $0.699 |
| Kimi K3 | 79.2 | 90.7 | 62.2% | $0.348 |
Fable 5 holds the overall crown by 2.9 points. Opus 5’s case is narrower and more specific: if the workload is agentic coding specifically, it outperforms the overall leaderboard leader at half the cost. For engineering teams deploying coding agents at scale, that is the more relevant comparison.