GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 586 -0.5%
INKL 531
CL-OP46 497
CL-OP48 490 -0.2%
← Back to feed

Claude Opus 5 Tops LiveBench Agentic Coding at 65.2% — Fable 5 Trails by 3 Points at Twice the Cost

The dominant pattern in frontier AI benchmarking has been Fable 5 holding the top line on most overall scores. LiveBench’s agentic coding dimension breaks that pattern. Claude Opus 5 Thinking Max Effort scores 65.2% on agentic coding, three points above Fable 5 Max Effort’s 62.2% and nine points above GPT-5.6 Sol’s 56.2%.

Across the full LiveBench leaderboard, Fable 5 leads at 83.0 overall. Opus 5 sits fourth at 80.1. But the models sort differently when the task is specifically agentic coding: Opus 5 moves to first, Kimi K3 ties Fable 5 at 62.2% in second, and Sol falls to fifth.

The Cost Math

The score difference is notable. The cost difference is sharper:

ModelAgentic CodingCost Per Task
Claude Opus 5 Max Effort65.2%$0.699
Claude Fable 5 Max Effort62.2%$1.439
Kimi K362.2%$0.348
GPT-5.5 xHigh54.0%$0.435
GPT-5.6 Sol Max Effort56.2%$0.438

Opus 5 beats Fable 5 on the metric and costs exactly half as much per successful task. Kimi K3 is the cheapest frontier-class option and ties Fable 5 at 62.2%; it loses to Opus 5 on score but undercuts it on cost at $0.348.

Why Agentic Coding Separates Here

LiveBench’s overall scores are dominated by reasoning, mathematics, and data analysis. These dimensions reward raw intelligence, where Fable 5’s training leads. Agentic coding requires a different profile: multi-step execution, tool use, context retention across long code sessions, and self-correction when initial approaches fail.

Opus 5 was positioned from its launch as a research and complex-task model. Its ARC-AGI-3 score of 30.2% — four times the previous best — reflects the same strength: sustained multi-step problem-solving that degrades slowly as session length increases. The LiveBench agentic coding result is independent confirmation of that design decision. On the dimension most closely resembling a real software engineering session, Opus 5 is the performance leader.

Artificial Analysis Cross-Check

The result aligns with Artificial Analysis data published this week. On the AA Intelligence Index, Opus 5 Adaptive Reasoning Max Effort scores 61, two points above GPT-5.6 Sol at 59. AA pricing data puts Opus 5 at $3.85 per million tokens blended versus Sol at $4.35, a 12% cost advantage. Time-to-first-token: Opus 5 at 65 seconds versus Sol at 143 seconds, a 2.2x response-time lead.

The two sources cross-reference cleanly. On AA Intelligence Index, Opus 5 beats Sol. On LiveBench agentic coding, Opus 5 beats Fable 5. On cost per successful task, Opus 5 undercuts Fable 5 by 51%.

Full LiveBench Top-5 Overview

ModelOverallReasoningAgentic CodingCost/Task
Fable 5 Max Effort83.089.762.2%$1.439
GPT-5.6 Sol Max Effort81.091.756.2%$0.438
GPT-5.5 xHigh80.289.754.0%$0.435
Claude Opus 5 Max Effort80.191.265.2%$0.699
Kimi K379.290.762.2%$0.348

Fable 5 holds the overall crown by 2.9 points. Opus 5’s case is narrower and more specific: if the workload is agentic coding specifically, it outperforms the overall leaderboard leader at half the cost. For engineering teams deploying coding agents at scale, that is the more relevant comparison.