GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Fable 5 Max Effort Costs $1.57 Per Task and Posts Agentic Coding's Weakest Score in the Top Tier at 46.9%

Claude Fable 5 now has a Max Effort entry on LiveBench. The headline number improves: 80.8 overall, up from 79.5 at xHigh. The agentic coding number goes the other direction.

At Max Effort, Fable 5 posts 46.9% on LiveBench’s Agentic Coding category — 3.8 points below its own 50.7% at xHigh effort. It ranks last among the current top-six models on the metric, behind GPT-5.4 Thinking xHigh (53.8%), GPT-5.5 Thinking xHigh (52.1%), Claude Opus 4.8 Thinking xHigh (56.1%), GPT-5.6 Sol Max Effort (65.6%), and GPT-5.6 Terra Max Effort (68.0%).

The cost doubles simultaneously. Fable 5 Max Effort costs $1.573 per successful task. The xHigh result came in at $0.811.

The Full Current Table

ModelOverallAgentic CodingLangReasoning$/task
GPT-5.6 Sol Max82.465.6%87.791.7$0.589
Fable 5 Max80.846.9%90.789.7$1.573
GPT-5.5 xHigh79.952.1%87.489.7$0.530
GPT-5.6 Terra Max79.868.0%82.990.6$0.497
Opus 4.8 Thinking xHigh78.956.1%81.489.7$0.688
GPT-5.4 Thinking xHigh78.053.8%82.688.1$0.387

Fable 5 Max leads on language (90.7%) and is competitive on reasoning (89.7%), tied with GPT-5.5 and Opus 4.8). Those categories justify its overall rank. The agentic coding result sits 18 points below the category leader and 19 below the runner-up, on a 100-point scale where the gap from first to fifth is 15 points.

Why Max Effort Harms Agentic Coding

LiveBench’s Agentic Coding category tests multi-step, tool-using, environment-interaction tasks — the same class of work that SWE-bench and tau2-bench measure. Extended thinking at Max Effort adds reasoning tokens before each action. For tasks requiring fluid, reactive tool-calling sequences, that additional deliberation appears to create overhead that degrades execution performance.

This pattern has appeared elsewhere. Anthropic’s own Claude Sonnet 5 benchmarks showed that extended thinking improved math and reasoning scores while leaving coding agent performance flat. The Max Effort LiveBench result for Fable 5 is the sharpest version of that dynamic yet measured: not flat, but negative, at a cost premium that makes the regression expensive.

The Terra Comparison

GPT-5.6 Terra Max Effort is the direct cost-tier alternative. It posts 68.0% on agentic coding — 21 points above Fable 5 Max — at $0.497 per task, 68% cheaper. Terra trails Fable 5 on language (82.9 vs 90.7) and instruction following (64.6 vs 75.8), but for teams running coding agents, the gap is 21 points in the wrong direction for Fable 5.

At 100 million successful agentic tasks per month, Fable 5 Max costs $157.3M annually versus Terra Max at $49.7M. The $107.6M difference buys a 21-point capability deficit on the category being priced.

What Has Not Changed

Fable 5 holds its position on Agent Arena, where it ranks first with a 12.94% composite on live user evaluations. Agent Arena uses pairwise human preference voting, not timed multi-step execution benchmarks. The LiveBench finding is specific to the Agentic Coding category as LiveBench defines it — it does not overturn Fable 5’s overall agentic standing.

The relevant implication is narrower: for automated pipelines running timed, multi-step coding agents at scale, Max Effort is the wrong mode for Fable 5. The model’s Max Effort budget is better applied to reasoning, writing, and language tasks where the improvement is real and the cost premium is more defensible.