LiveBench July 2026: Claude Opus 4.8 Leads Agentic Coding at 56.1% — GPT-5.5 Wins Overall, Fable 5 Costs 53% More Per Task
The July 2026 LiveBench standings surface a divergence that headline ELO rankings obscure: the model ranked first overall is not the best model for agentic coding tasks. Claude Opus 4.8 Thinking leads the Agentic Coding category at 56.1% — ahead of GPT-5.5 (52.1%) and Claude Fable 5 (50.7%) — despite ranking third on the overall composite.
Current Rankings
| Model | Overall | Agentic Coding | Math | Language | $/task |
|---|---|---|---|---|---|
| GPT-5.5 Thinking xHigh | 79.9 | 52.1 | 95.9 | 87.4 | $0.530 |
| Claude Fable 5 xHigh | 79.5 | 50.7 | 95.7 | 89.5 | $0.811 |
| Claude Opus 4.8 Thinking xHigh | 78.9 | 56.1 | 95.3 | 81.4 | $0.688 |
| GPT-5.4 Thinking xHigh | 78.0 | 53.8 | 94.1 | 82.6 | $0.387 |
| Gemini 3.1 Pro Preview High | 77.1 | 45.4 | 91.0 | 85.4 | $0.262 |
| Claude Opus 4.7 Thinking xHigh | 76.5 | 50.7 | 92.9 | 77.9 | $0.528 |
LiveBench evaluates on continuously refreshed tasks across Reasoning, Coding, Agentic Coding, Mathematics, Data Analysis, Language, and Instruction Following. The cost column is cost per successful task — not per million tokens.
The Agentic Coding Split
The Agentic Coding category is the most significant divergence. Claude Opus 4.8 leads at 56.1%, ahead of GPT-5.4 (53.8%), GPT-5.5 (52.1%), and Claude Fable 5 (50.7%). Gemini 3.1 Pro trails by more than 10 points at 45.4%.
Fable 5 outranks Opus 4.8 on the overall composite by 0.6 points but trails by 5.4 points on the category designed to test multi-step, tool-using agent workflows. That gap matters in production: SWE-bench and tau2-bench also scored Opus 4.8’s Thinking variant at or above the ceiling. LiveBench’s Agentic Coding results are consistent with those findings and add a third independent signal in the same direction.
GPT-5.4, typically one tier below the current flagship, posts 53.8% on Agentic Coding — second place — at $0.387 per successful task, the second-cheapest option in the top four.
Math Is a Dead Heat
At the top three positions, mathematics scores span 0.6 points: GPT-5.5 at 95.9, Fable 5 at 95.7, Opus 4.8 at 95.3. For purely reasoning-intensive tasks, all three frontier models are functionally equivalent. Differentiation on math capability at the frontier is exhausted.
Language Is Fable 5’s Edge
Fable 5 leads language at 89.5%, 2.1 points above GPT-5.5 (87.4%) and 8.1 points above Opus 4.8 (81.4%). The 8-point gap on language versus the near-identical math and similar overall scores suggests Fable 5’s strongest domain advantage is prose quality, summarisation, and writing coherence — not agentic execution.
Gemini 3.1 Pro scores 85.4 on language, placing third above both OpenAI and Anthropic’s non-language-specialist variants. At $0.262 per successful task, it is the only model to hit the 85+ language threshold under $0.30 per task.
Cost Per Successful Task
The cost-per-task comparison changes the model selection calculus:
- GPT-5.5 at $0.530 is 23% cheaper per successful task than Opus 4.8 ($0.688)
- Fable 5 at $0.811 costs 53% more per task than GPT-5.5 and 18% more than Opus 4.8
- GPT-5.4 at $0.387 costs 32% less than GPT-5.5 with a 1.9-point overall penalty
- Gemini 3.1 Pro at $0.262 is the price-performance leader for organisations willing to accept a 2.8-point overall gap versus GPT-5.5
At 100 million successful agent tasks per month — a realistic scale for large enterprise deployments — the spread between Fable 5 ($81.1M) and GPT-5.5 ($53.0M) is $28.1M annually. The Gemini option ($26.2M) nearly halves the Fable 5 bill.
Selection Framework
LiveBench’s category breakdown points toward a simple framework:
- Agentic coding pipelines: Claude Opus 4.8 Thinking or GPT-5.4 Thinking lead on the metric that counts
- Language-heavy applications (legal, editorial, customer communications): Claude Fable 5 leads by 2-8 points depending on competitor
- Cost-sensitive at-scale deployments: GPT-5.4 Thinking or Gemini 3.1 Pro depending on whether agentic or language tasks dominate
- Peak overall performance, budget unconstrained: GPT-5.5 Thinking xHigh
The “best model” is a routing problem, not a model selection problem. The LiveBench data makes the routing inputs explicit.