GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
← Back to feed

Claude Opus 5 Enters LiveBench at #3 Overall — Beats Fable 5 on Agentic Coding 61.3% to 46.9%

Claude Opus 5 Thinking xHigh has entered the LiveBench leaderboard at 80.3 overall — third place, behind GPT-5.6 Sol Max Effort (82.4) and Claude Fable 5 Max Effort (80.8). It is the first Anthropic model outside the Fable/Mythos line to clear 80 on the overall composite.

The more significant result is the Agentic Coding category.

Agentic Coding: Opus 5 vs Fable 5

ModelOverallAgentic Coding$/task
GPT-5.6 Sol Max Effort82.465.6$0.589
Claude Fable 5 Max Effort80.846.9$1.573
Claude Opus 5 Thinking xHigh80.361.3$0.487
GPT-5.5 Thinking xHigh79.952.1$0.530
GPT-5.6 Terra Max Effort79.868.0$0.497
Claude Opus 4.8 Thinking xHigh78.956.1$0.688

Fable 5 posts 46.9% on LiveBench Agentic Coding. Opus 5 posts 61.3% — a 14.4-point gap on the category designed to test multi-step, tool-using agent workflows.

That gap is not close. Among models above 79 overall, only GPT-5.6 Sol (65.6%) and GPT-5.6 Terra (68.0%) post higher agentic coding scores than Opus 5.

Fable 5’s 46.9% on agentic coding has been a consistent finding. It leads on language (90.7%) and mathematics (96.0%), but agentic execution is its documented weakness. Opus 5 flips that relationship: 61.3% agentic coding at 94.8% mathematics — within 1.2 points of Fable 5 on math, 14.4 ahead on autonomous work.

Cost

Opus 5 at $0.487 per successful task is 69% cheaper than Fable 5 at $1.573. It is also cheaper than GPT-5.5 ($0.530) and GPT-5.6 Sol ($0.589).

The only model in the top-six bracket cheaper than Opus 5 is GPT-5.6 Terra at $0.497 — a difference of $0.01 per task, with Terra leading agentic coding by 6.7 points.

Context

The July 24 Opus 5 launch claimed improvements in Frontier-Bench and ARC-AGI-3 relative to Fable 5 and Opus 4.8. LiveBench’s Agentic Coding category adds a third independent signal consistent with those results: Opus 5 closes the frontier agentic gap that made Fable 5 the default choice for complex coding agent deployments.

Fable 5 retains the lead on language and mathematics. If the task requires sustained multi-step autonomous execution, the numbers now point somewhere else.