GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Claude Fable 5 Takes #1 on Artificial Analysis Intelligence Index at 64.9 — Anthropic Holds Both Top Spots

Claude Fable 5 scores 64.9 on the Artificial Analysis Intelligence Index, the 10-benchmark composite that combines reasoning, knowledge, mathematics, coding, and agentic capability. That puts it above Claude Opus 4.8 (61.4, #2) and GPT-5.5 (~60, #3), making Anthropic the first lab to hold both the top two positions simultaneously.

The gap to the next non-Anthropic model is approximately 4–5 points. The previous record was Opus 4.8’s 61.4, set in late May. Fable 5’s 64.9 is the single largest jump between a #1 position and its predecessor in the index’s recent history.

What the Index Measures

AA Intelligence Index v4.0 incorporates 10 evaluations: GDPval-AA (real-world work tasks), Tau2-bench Telecom (tool use), Terminal-Bench Hard (agentic coding), SciCode, AA-LCR, AA-Omniscience (knowledge and hallucination), IFBench (instruction following), Humanity’s Last Exam, GPQA Diamond, and CritPt. Fable 5 leads on five of the ten.

Key Benchmark Numbers

Humanity’s Last Exam (HLE): 53%

Fable 5 scores 53% on HLE, the crowdsourced doctoral-level benchmark. The next-best model is Claude Opus 4.8 (max), which sits more than 7 points behind. There is no other model within 10 points of Fable 5 on this benchmark.

The asterisk: Fable 5 triggers safety guardrails on 9% of HLE tasks, falling back to Opus 4.8 mid-trajectory. The AA team ran the evaluation including those fallbacks. Total cost to run HLE with Fable 5 came to approximately $2.2k — the highest of any model they have evaluated — reflecting the per-token cost of Fable 5 ($10/$50 input/output per million) and the overhead of partial Opus 4.8 sessions.

AA-Omniscience: 40

AA’s knowledge and hallucination benchmark measures what a model knows and how often it confabulates. Fable 5 scores 40, up 7 points over the previous leader, Gemini 3.1 Pro Preview. The AA team notes this improvement is driven primarily by higher accuracy (more questions answered correctly), not by lower hallucination rates. That pattern — more accuracy, not less hallucination — aligns with what they observe in open-weight models as parameter count increases, which they interpret as a possible signal that Fable 5 is larger than any previous publicly released Anthropic model.

GDPval-AA Elo: 1932

The GDPval-AA Elo measures real-world work task performance. Fable 5 comes in at 1932, up from Opus 4.8’s 1890 — a 42-point jump. The previous generation leader was GPT-5.5, which sits at 1769.

GPQA Diamond: 92.6%

Graduate-level scientific reasoning. Fable 5 scores 92.6%, consistent with leading performance across scientific domains.

Agentic evaluations: frontier across all three

The AA team places Fable 5 at the frontier on all three agentic benchmarks in the index — GDPval-AA, Terminal-Bench Hard, and Tau2-bench Telecom — without publishing specific per-benchmark numbers for those evaluations separately. The combined effect is the highest composite score the index has produced.

The Fallback Mechanics at Scale

Fable 5 uses the same underlying model as Mythos 5, wrapped in safety classifiers. Fallback to Opus 4.8 occurs on fewer than 5% of sessions in Anthropic’s own reporting, but the AA team recorded fallback routing in approximately 8% of tasks across the full Intelligence Index — concentrated in scientific questions from GPQA, AA-Omniscience, and HLE.

That gap (5% vs 8%) reflects the distribution of Index queries, which over-weights hard scientific domains relative to a typical production API mix. The practical implication: applications in research, biology, and chemistry will see higher fallback rates than general coding or knowledge work.

Competitive Positioning

ModelAA Intelligence IndexHLEGDPval-AA Elo
Claude Fable 564.953%1932
Claude Opus 4.8 (max)61.4~46%1890
GPT-5.5~60~46%1769
Gemini 3.1 Pro Preview~57~43%1314

The 3.5-point gap between Fable 5 and Opus 4.8 is meaningful. Opus 4.8 was itself the record holder two weeks ago. The same gap between Opus 4.8 and GPT-5.5 at #3 gives a sense of scale: Fable 5 is roughly two full model generations ahead of GPT-5.5 on this composite.

Pricing and Access

Fable 5 is priced at $10 per million input tokens and $50 per million output tokens — exactly double Opus 4.8 ($5/$25). Access through Pro, Max, Team, and Enterprise plans is included at no extra cost through June 22. From June 23, usage credits are required, with restoration to standard subscriptions planned once capacity permits.

API identifier: claude-fable-5. Mythos 5, the same underlying model without safety classifiers in restricted domains, remains limited to Project Glasswing partners.