GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Fable 5 Returns Worldwide as GPT-5.6 Sol Stays Government-Only: A Benchmark Split With an Access Wall

Two weeks of regulatory turbulence produced an unusual result for enterprise AI buyers: the two best models in the world are now on completely different release tracks, and their benchmark leads run in opposite directions.

What Happened

Anthropic’s Claude Fable 5 went back online globally on July 1, two days after the Commerce Department cleared both Fable 5 and Mythos 5 following a 19-day suspension. The clearance came after Amazon researchers surfaced a jailbreak capable of producing functional exploit code on June 12. Anthropic is offering a usage bonus for paid subscribers through July 7.

GPT-5.6 Sol, OpenAI’s current flagship, has not moved. As of July 3, access remains limited to approximately 20 government-vetted organisations via the API and Codex. OpenAI delayed the broader launch at the White House’s request while national-security cybersecurity capability reviews run to completion. General availability is described as arriving “in the coming weeks.” No date has been set.

The Benchmark Split

The two models lead on different tests, and the gaps are not small.

Terminal-Bench 2.1: Sol 88.8%. Sol Ultra mode, which spins up coordinated subagents for complex work, reaches 91.9% — the highest published mark on that leaderboard. Fable 5’s published Terminal-Bench 2.1 figure is 84.3%.

SWE-Bench Pro: Fable 5 80.3%. GPT-5.5 trails at 58.6%, a 21.7-point gap. OpenAI has published no SWE-Bench Pro figure for GPT-5.6 at all. The benchmark most evaluators treat as decisive for autonomous software work — end-to-end fixes on real GitHub issues, not command-line execution — belongs entirely to Anthropic at the public frontier tier.

The split matters because the two benchmarks measure different things. Terminal-Bench 2.1 measures command-line agentic coordination: planning, iteration, tool orchestration. SWE-Bench Pro measures whether a model can actually close a real-world software bug end to end. Enterprises choosing a coding agent are choosing which capability profile they need, not just which benchmark score is higher.

The Access Problem

The deeper issue for buyers is that the competitive landscape is now partly hypothetical. Fable 5 is purchasable today at its standard pricing. Sol is not available to most organisations at any price.

When Sol does reach general availability, the benchmark comparison will reset again — OpenAI has confirmed Sol includes an Ultra mode with parallel subagents, and the company typically publishes SWE-Bench Verified results, not Pro, which makes direct comparison with Fable 5 difficult by design.

For now, the question “which is better” has a concrete practical answer: Fable 5 is available, and it leads the benchmark that matters most for code-centric autonomous work by 21 points over any model you can actually buy.

Key Numbers

BenchmarkFable 5GPT-5.6 SolGPT-5.5
SWE-Bench Pro80.3%Not published58.6%
Terminal-Bench 2.184.3%88.8% (91.9% Ultra)~78.2%
AccessGlobal (July 1)~20 gov-vetted orgsGenerally available

The government’s decision to stage both releases simultaneously has created the cleanest natural experiment the frontier has produced: same class of model, same target use case, different benchmarks, different keys.