GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

LiveBench July 2026: Claude Opus 4.8 Leads Agentic Coding at 56.1% — GPT-5.5 Wins Overall, Fable 5 Costs 53% More Per Task

The July 2026 LiveBench standings surface a divergence that headline ELO rankings obscure: the model ranked first overall is not the best model for agentic coding tasks. Claude Opus 4.8 Thinking leads the Agentic Coding category at 56.1% — ahead of GPT-5.5 (52.1%) and Claude Fable 5 (50.7%) — despite ranking third on the overall composite.

Current Rankings

ModelOverallAgentic CodingMathLanguage$/task
GPT-5.5 Thinking xHigh79.952.195.987.4$0.530
Claude Fable 5 xHigh79.550.795.789.5$0.811
Claude Opus 4.8 Thinking xHigh78.956.195.381.4$0.688
GPT-5.4 Thinking xHigh78.053.894.182.6$0.387
Gemini 3.1 Pro Preview High77.145.491.085.4$0.262
Claude Opus 4.7 Thinking xHigh76.550.792.977.9$0.528

LiveBench evaluates on continuously refreshed tasks across Reasoning, Coding, Agentic Coding, Mathematics, Data Analysis, Language, and Instruction Following. The cost column is cost per successful task — not per million tokens.

The Agentic Coding Split

The Agentic Coding category is the most significant divergence. Claude Opus 4.8 leads at 56.1%, ahead of GPT-5.4 (53.8%), GPT-5.5 (52.1%), and Claude Fable 5 (50.7%). Gemini 3.1 Pro trails by more than 10 points at 45.4%.

Fable 5 outranks Opus 4.8 on the overall composite by 0.6 points but trails by 5.4 points on the category designed to test multi-step, tool-using agent workflows. That gap matters in production: SWE-bench and tau2-bench also scored Opus 4.8’s Thinking variant at or above the ceiling. LiveBench’s Agentic Coding results are consistent with those findings and add a third independent signal in the same direction.

GPT-5.4, typically one tier below the current flagship, posts 53.8% on Agentic Coding — second place — at $0.387 per successful task, the second-cheapest option in the top four.

Math Is a Dead Heat

At the top three positions, mathematics scores span 0.6 points: GPT-5.5 at 95.9, Fable 5 at 95.7, Opus 4.8 at 95.3. For purely reasoning-intensive tasks, all three frontier models are functionally equivalent. Differentiation on math capability at the frontier is exhausted.

Language Is Fable 5’s Edge

Fable 5 leads language at 89.5%, 2.1 points above GPT-5.5 (87.4%) and 8.1 points above Opus 4.8 (81.4%). The 8-point gap on language versus the near-identical math and similar overall scores suggests Fable 5’s strongest domain advantage is prose quality, summarisation, and writing coherence — not agentic execution.

Gemini 3.1 Pro scores 85.4 on language, placing third above both OpenAI and Anthropic’s non-language-specialist variants. At $0.262 per successful task, it is the only model to hit the 85+ language threshold under $0.30 per task.

Cost Per Successful Task

The cost-per-task comparison changes the model selection calculus:

  • GPT-5.5 at $0.530 is 23% cheaper per successful task than Opus 4.8 ($0.688)
  • Fable 5 at $0.811 costs 53% more per task than GPT-5.5 and 18% more than Opus 4.8
  • GPT-5.4 at $0.387 costs 32% less than GPT-5.5 with a 1.9-point overall penalty
  • Gemini 3.1 Pro at $0.262 is the price-performance leader for organisations willing to accept a 2.8-point overall gap versus GPT-5.5

At 100 million successful agent tasks per month — a realistic scale for large enterprise deployments — the spread between Fable 5 ($81.1M) and GPT-5.5 ($53.0M) is $28.1M annually. The Gemini option ($26.2M) nearly halves the Fable 5 bill.

Selection Framework

LiveBench’s category breakdown points toward a simple framework:

  • Agentic coding pipelines: Claude Opus 4.8 Thinking or GPT-5.4 Thinking lead on the metric that counts
  • Language-heavy applications (legal, editorial, customer communications): Claude Fable 5 leads by 2-8 points depending on competitor
  • Cost-sensitive at-scale deployments: GPT-5.4 Thinking or Gemini 3.1 Pro depending on whether agentic or language tasks dominate
  • Peak overall performance, budget unconstrained: GPT-5.5 Thinking xHigh

The “best model” is a routing problem, not a model selection problem. The LiveBench data makes the routing inputs explicit.