GPT-5.6 Sol Tops LiveBench at 82.4 and Leads AA Coding Agent Index at 80 — at One-Third the Cost of Fable 5
GPT-5.6 Sol has cleared the bar set by its own launch benchmarks. Independent evaluations from LiveBench and Artificial Analysis, published this week, place Sol at the head of the overall frontier leaderboard and atop the coding agent rankings — at a price that undercuts the previous quality leader by 70%.
LiveBench: Sol at 82.4, the Highest Overall Score Recorded
LiveBench’s July 10 update puts GPT-5.6 Sol Max Effort at 82.4 overall — 2.5 points above GPT-5.5 Thinking xHigh (79.9), which held the top spot in the July 9 snapshot that Stack Futures covered. GPT-5.6 Terra sits at 79.8 and Claude Fable 5 xHigh at 79.5. Claude Opus 4.8 Thinking falls to fifth at 78.9.
Category-level scores for Sol (max effort):
| Category | GPT-5.6 Sol | GPT-5.5 xHigh | Fable 5 xHigh | Opus 4.8 Thinking |
|---|---|---|---|---|
| Overall | 82.4 | 79.9 | 79.5 | 78.9 |
| Reasoning | 91.7 | 89.7 | 87.7 | 89.7 |
| Coding | 83.9 | 82.1 | 82.5 | 79.3 |
| Agentic Coding | 65.6 | 52.1 | 50.7 | 56.1 |
| Mathematics | 96.2 | 95.9 | 95.7 | 95.3 |
| Data Analysis | 79.8 | 81.6 | 78.7 | 78.3 |
| Language | 87.7 | 87.4 | 89.5 | 81.4 |
| Instruction Following | 71.8 | 70.7 | 72.0 | 72.4 |
| Cost per task | $0.589 | $0.530 | $0.811 | $0.688 |
Sol leads in five of eight categories. Fable 5 retains the language and instruction-following edges but trails Sol by 3 points overall.
One anomaly worth noting: GPT-5.6 Terra leads Agentic Coding at 68.0%, above Sol’s 65.6. Terra costs $0.497 per task versus Sol’s $0.589. For teams running coding agents specifically, Terra offers a better score at lower cost — a distinction that aggregate rankings obscure.
Artificial Analysis: One Point Below Fable 5, One-Third the Price
Artificial Analysis published its evaluation of the full GPT-5.6 line on July 9. GPT-5.6 Sol (max) scores 59 on the AA Intelligence Index, one point below Claude Fable 5. Terra lands at 55, Luna at 51.
On the AA Coding Agent Index, Sol (max) leads the entire field at 80 points — ahead of Fable 5, Opus 4.8, and GPT-5.5.
The cost contrast is the sharpest data point in the report. Sol runs at $1.04 per Intelligence Index task; Fable 5 costs roughly three times that figure. Terra at $0.55 and Luna at $0.21 extend the efficiency frontier downward. Luna, according to AA, matches GLM-5.2 and Gemini 3.5 Flash at a lower per-task cost.
What the Data Establishes
Three takeaways from a week of third-party measurement:
-
Sol is the new overall benchmark leader. LiveBench overall and the AA Coding Agent Index both put it at the top, one week after launch. The pre-release government-gate controversy has not altered the model’s measured capability.
-
Fable 5 holds a narrow quality edge on AA Intelligence Index (60 vs 59) but at a large cost premium. Teams benchmarking raw reasoning tasks will continue to favour Fable 5. Teams running large volumes of agent tasks will find Sol financially compelling.
-
Terra is the underreported story. 68.0% on LiveBench Agentic Coding — the highest in the table — with a lower per-task cost than Sol. For agentic workloads specifically, Terra is the current value-performance optimum in the GPT-5.6 lineup.
The previous LiveBench #1 position (GPT-5.5 at 79.9) held for roughly ten days following the July 9 update before Sol’s entry moved the marker. Claude Fable 5’s overall score has been stable at 79.5 since its June launch.