GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

GPT-5.6 Sol Tops LiveBench at 82.4 and Leads AA Coding Agent Index at 80 — at One-Third the Cost of Fable 5

GPT-5.6 Sol has cleared the bar set by its own launch benchmarks. Independent evaluations from LiveBench and Artificial Analysis, published this week, place Sol at the head of the overall frontier leaderboard and atop the coding agent rankings — at a price that undercuts the previous quality leader by 70%.

LiveBench: Sol at 82.4, the Highest Overall Score Recorded

LiveBench’s July 10 update puts GPT-5.6 Sol Max Effort at 82.4 overall — 2.5 points above GPT-5.5 Thinking xHigh (79.9), which held the top spot in the July 9 snapshot that Stack Futures covered. GPT-5.6 Terra sits at 79.8 and Claude Fable 5 xHigh at 79.5. Claude Opus 4.8 Thinking falls to fifth at 78.9.

Category-level scores for Sol (max effort):

CategoryGPT-5.6 SolGPT-5.5 xHighFable 5 xHighOpus 4.8 Thinking
Overall82.479.979.578.9
Reasoning91.789.787.789.7
Coding83.982.182.579.3
Agentic Coding65.652.150.756.1
Mathematics96.295.995.795.3
Data Analysis79.881.678.778.3
Language87.787.489.581.4
Instruction Following71.870.772.072.4
Cost per task$0.589$0.530$0.811$0.688

Sol leads in five of eight categories. Fable 5 retains the language and instruction-following edges but trails Sol by 3 points overall.

One anomaly worth noting: GPT-5.6 Terra leads Agentic Coding at 68.0%, above Sol’s 65.6. Terra costs $0.497 per task versus Sol’s $0.589. For teams running coding agents specifically, Terra offers a better score at lower cost — a distinction that aggregate rankings obscure.

Artificial Analysis: One Point Below Fable 5, One-Third the Price

Artificial Analysis published its evaluation of the full GPT-5.6 line on July 9. GPT-5.6 Sol (max) scores 59 on the AA Intelligence Index, one point below Claude Fable 5. Terra lands at 55, Luna at 51.

On the AA Coding Agent Index, Sol (max) leads the entire field at 80 points — ahead of Fable 5, Opus 4.8, and GPT-5.5.

The cost contrast is the sharpest data point in the report. Sol runs at $1.04 per Intelligence Index task; Fable 5 costs roughly three times that figure. Terra at $0.55 and Luna at $0.21 extend the efficiency frontier downward. Luna, according to AA, matches GLM-5.2 and Gemini 3.5 Flash at a lower per-task cost.

What the Data Establishes

Three takeaways from a week of third-party measurement:

  1. Sol is the new overall benchmark leader. LiveBench overall and the AA Coding Agent Index both put it at the top, one week after launch. The pre-release government-gate controversy has not altered the model’s measured capability.

  2. Fable 5 holds a narrow quality edge on AA Intelligence Index (60 vs 59) but at a large cost premium. Teams benchmarking raw reasoning tasks will continue to favour Fable 5. Teams running large volumes of agent tasks will find Sol financially compelling.

  3. Terra is the underreported story. 68.0% on LiveBench Agentic Coding — the highest in the table — with a lower per-task cost than Sol. For agentic workloads specifically, Terra is the current value-performance optimum in the GPT-5.6 lineup.

The previous LiveBench #1 position (GPT-5.5 at 79.9) held for roughly ten days following the July 9 update before Sol’s entry moved the marker. Claude Fable 5’s overall score has been stable at 79.5 since its June launch.