GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

LiveBench Cost-Per-Task Metric Shows 9x Gap: Gemini 3.7 Flash High Beats Claude Fable 5 on Value

The August 2026 LiveBench standings reveal something the headline numbers miss: the cost per successful task varies by 9x across frontier models, while the quality gap spans only 4.2 points.

Claude Fable 5 Max Effort leads at 83.0 overall. Its cost: $1.439 per successful task. Gemini 3.7 Flash High sits at 78.8 and charges $0.157 per task. Nine times cheaper for four fewer quality points.

The Full Picture

ModelOverallAgentic CodingCost/Task
Claude Fable 5 Max Effort83.062.2$1.439
GPT-5.6 Sol Max Effort81.056.2$0.515
GPT-5.5 Thinking xHigh80.254.0$0.435
Claude 5 Opus Thinking Max Effort80.165.2$0.699
Kimi K3 open79.262.2$0.348
Gemini 3.7 Flash High78.858.3$0.157

What the Numbers Say

The quality compression is real. Six models sit within 4.2 points of each other. The economics are not compressed at all.

GPT-5.6 Sol at $0.515/task represents the clearest mid-tier value: 2 points behind Fable 5 at 36% of the cost. For most enterprise applications, that trade is rational.

The open-weight outlier is Kimi K3. At $0.348/task with 79.2 overall — including 62.2 on Agentic Coding, matching Fable 5’s score on that dimension — it is the most cost-competitive open model in the current LiveBench top ten. Teams self-hosting or running at scale have a credible frontier-adjacent option without API lock-in.

Gemini 3.7 Flash High is in a separate cost category. Its $0.157/task is roughly half of Kimi K3 and less than a tenth of Fable 5. Its agentic coding score (58.3) trails Fable 5 by 4 points. For workloads running thousands of tasks per day, that cost delta compounds fast. A pipeline making 10,000 daily calls pays $1,570 with Fable 5 and $157 with Flash High — the $1,400 daily savings funds a meaningful portion of engineering headcount.

The Agentic Coding Exception

The one dimension where leaderboard rank diverges sharply from overall scores: Agentic Coding.

Claude 5 Opus Thinking Max Effort leads at 65.2% — beating both Fable 5 (62.2) and GPT-5.6 Sol (56.2), despite ranking fourth overall. Its $0.699/task is half of Fable 5. Teams optimising specifically for agentic tasks — code generation, debugging loops, multi-step tool use — are better served by Opus 5 Thinking than by Fable 5 at twice the price.

GPT-5.6 Sol’s 56.2 on Agentic Coding is the weakest score in the top six, despite posting the second-best overall. That gap is unusually wide for a frontier model and may reflect optimisation for chat-style benchmarks over multi-turn tool calling.

Key Numbers

  • Cost range: $0.157 (Gemini 3.7 Flash High) to $1.439 (Claude Fable 5) — 9.2x spread
  • Quality range: 78.8 to 83.0 — 4.2 point spread
  • Best overall value per quality point: Gemini 3.7 Flash High
  • Best agentic coding value: Claude 5 Opus Thinking at $0.699/task, 65.2% agentic coding
  • Open-weight value leader: Kimi K3 at $0.348/task, 79.2 overall, matching Fable 5 on agentic coding
  • GPT-5.6 Sol agentic coding underperformance: 56.2 vs 81.0 overall — 24.8 point drop, widest in the top six