LiveBench Cost-Per-Task Metric Shows 9x Gap: Gemini 3.7 Flash High Beats Claude Fable 5 on Value
The August 2026 LiveBench standings reveal something the headline numbers miss: the cost per successful task varies by 9x across frontier models, while the quality gap spans only 4.2 points.
Claude Fable 5 Max Effort leads at 83.0 overall. Its cost: $1.439 per successful task. Gemini 3.7 Flash High sits at 78.8 and charges $0.157 per task. Nine times cheaper for four fewer quality points.
The Full Picture
| Model | Overall | Agentic Coding | Cost/Task |
|---|---|---|---|
| Claude Fable 5 Max Effort | 83.0 | 62.2 | $1.439 |
| GPT-5.6 Sol Max Effort | 81.0 | 56.2 | $0.515 |
| GPT-5.5 Thinking xHigh | 80.2 | 54.0 | $0.435 |
| Claude 5 Opus Thinking Max Effort | 80.1 | 65.2 | $0.699 |
| Kimi K3 open | 79.2 | 62.2 | $0.348 |
| Gemini 3.7 Flash High | 78.8 | 58.3 | $0.157 |
What the Numbers Say
The quality compression is real. Six models sit within 4.2 points of each other. The economics are not compressed at all.
GPT-5.6 Sol at $0.515/task represents the clearest mid-tier value: 2 points behind Fable 5 at 36% of the cost. For most enterprise applications, that trade is rational.
The open-weight outlier is Kimi K3. At $0.348/task with 79.2 overall — including 62.2 on Agentic Coding, matching Fable 5’s score on that dimension — it is the most cost-competitive open model in the current LiveBench top ten. Teams self-hosting or running at scale have a credible frontier-adjacent option without API lock-in.
Gemini 3.7 Flash High is in a separate cost category. Its $0.157/task is roughly half of Kimi K3 and less than a tenth of Fable 5. Its agentic coding score (58.3) trails Fable 5 by 4 points. For workloads running thousands of tasks per day, that cost delta compounds fast. A pipeline making 10,000 daily calls pays $1,570 with Fable 5 and $157 with Flash High — the $1,400 daily savings funds a meaningful portion of engineering headcount.
The Agentic Coding Exception
The one dimension where leaderboard rank diverges sharply from overall scores: Agentic Coding.
Claude 5 Opus Thinking Max Effort leads at 65.2% — beating both Fable 5 (62.2) and GPT-5.6 Sol (56.2), despite ranking fourth overall. Its $0.699/task is half of Fable 5. Teams optimising specifically for agentic tasks — code generation, debugging loops, multi-step tool use — are better served by Opus 5 Thinking than by Fable 5 at twice the price.
GPT-5.6 Sol’s 56.2 on Agentic Coding is the weakest score in the top six, despite posting the second-best overall. That gap is unusually wide for a frontier model and may reflect optimisation for chat-style benchmarks over multi-turn tool calling.
Key Numbers
- Cost range: $0.157 (Gemini 3.7 Flash High) to $1.439 (Claude Fable 5) — 9.2x spread
- Quality range: 78.8 to 83.0 — 4.2 point spread
- Best overall value per quality point: Gemini 3.7 Flash High
- Best agentic coding value: Claude 5 Opus Thinking at $0.699/task, 65.2% agentic coding
- Open-weight value leader: Kimi K3 at $0.348/task, 79.2 overall, matching Fable 5 on agentic coding
- GPT-5.6 Sol agentic coding underperformance: 56.2 vs 81.0 overall — 24.8 point drop, widest in the top six