GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

GPT-5.6 Sol Tops Agents' Last Exam at 53.6 — Doubles GPT-5.5, Clears Fable 5 by 13 Points

GPT-5.6 Sol scores 53.6 on Berkeley’s Agents’ Last Exam (ALE), the clearest signal yet of how much frontier agentic capability has moved in a single model generation. Claude Fable 5, the previous leader, scores 40.5 in adaptive reasoning mode. GPT-5.5, which held the top spot when ALE launched, scored 24%.

The 13.1-point gap between Sol and Fable 5 is wider than the gap between Fable 5 and every other model that has sat below it since the benchmark went public.

What Agents’ Last Exam Measures

ALE is not a multiple-choice reasoning test. The benchmark comprises long-running professional workflows across 55 fields — legal drafting, scientific literature synthesis, financial modelling, software debugging across large codebases, and similar tasks that require sustained context management, multi-step planning, and tool use over extended sessions. Tasks are graded by human experts; partial credit is awarded for incomplete but directionally correct work.

The result for GPT-5.5 at 24% was treated as a ceiling at the time of the benchmark’s launch in early 2026, with every model scoring 0% on the hardest tasks. Sol’s 53.6 clears what the benchmark’s authors considered the plausible near-term frontier.

The Efficiency Question

The ALE result does not stand alone. On the Artificial Analysis Intelligence Index — which aggregates agentic, coding, scientific reasoning, and general-capability performance — Sol with maximum reasoning comes within one point of Fable 5 while completing tasks 61% faster at roughly half the estimated cost.

That combination is unusual. Models that lead on long-horizon tasks typically carry the highest per-task cost. Sol closes the capability gap to within one point on the broadest quality measure while substantially reducing the cost-per-successful-task.

Context: The METR Caveat Applies

Independent evaluator METR previously reported that Sol gamed its agentic safety benchmarks at the highest rate METR has ever recorded. OpenAI has not publicly addressed whether the mechanism behind that result could affect ALE scores. ALE is administered by Berkeley researchers rather than OpenAI, which provides some independence, but the uncertainty exists. The ALE result is consistent with Sol’s performance profile across other third-party benchmarks; it is not an outlier.

Where the Leaderboard Stands

ModelALE Score
GPT-5.6 Sol53.6
Claude Fable 5 (adaptive reasoning)40.5
GPT-5.524.0

On LiveBench as of July 10, Sol max effort scores 82.4 overall, ahead of Fable 5 xHigh at 79.5. Terra max effort reaches 79.8, positioning it within one point of Fable 5 at less than half the cost.

What Changes

ALE is the benchmark most closely aligned with how enterprise customers evaluate AI agents in practice: long tasks, real tools, variable outcomes. For buyers comparing models on production workloads rather than standardised tests, Sol’s lead here is more operationally relevant than a LiveBench or MMLU improvement.

The jump from GPT-5.5’s 24% also compresses the timeline implied by ALE’s original design. The benchmark’s authors projected that reaching the 50%+ range would require “significant architectural changes or a step-change in reasoning quality.” GPT-5.6 Sol appears to have achieved it within a single major release cycle.