Fable 5 Takes Humanity's Last Exam at 53.3% — Field Average Is 12.6%
Humanity’s Last Exam, the expert-level reasoning benchmark built from PhD-level questions across mathematics, science, and professional domains, shows Claude Fable 5 leading at 53.3% as of the July 1, 2026 update. That score is 7.6 points ahead of Claude Opus 4.8 in second place at 45.7%, and 8.6 points ahead of Gemini 3.1 Pro Preview at 44.7%.
Of the 248 models currently evaluated on HLE, the mean score is 12.6% with a standard deviation of 10.8 points.
The Numbers
| Model | HLE Score |
|---|---|
| Claude Fable 5 | 53.3% |
| Claude Opus 4.8 | 45.7% |
| Gemini 3.1 Pro Preview | 44.7% |
| Field average (248 models) | 12.6% |
Anthropic holds the top two spots. The gap between Fable 5 and the second-place model from a different lab — Gemini 3.1 Pro Preview — is 8.6 points. The gap between Fable 5 and the field average is 40.7 points.
What Humanity’s Last Exam Tests
HLE was developed as a benchmark that graduate-level LLMs could not trivially saturate. Questions come from across advanced mathematics, physics, chemistry, biology, law, economics, and professional certification exams — many contributed by domain experts and professors. The design prioritizes questions where correct answers require genuine knowledge and multi-step reasoning rather than pattern matching against training data.
The benchmark appeared as a Spotlight Paper at ICLR 2025. It now tracks 248 models.
What the Gap Means
The 53.3% figure for Fable 5 represents a meaningful threshold. Most frontier models from the past 12 months — including GPT-5.5 and the open-weight field — cluster between 20% and 40% on this benchmark. A score above 50% requires consistently answering correctly across diverse expert domains, not just in the model’s strongest areas.
The 12.6% average across 248 models illustrates how hard the benchmark is. The majority of deployed models, including capable production systems, score below 15%. Fable 5 at 53.3% is operating in a regime that was not achievable by any public model at HLE’s launch.
The Opus 4.8 result (45.7%) is also notable: it shows the Mythos-class weights, even in their Opus-tier configuration, sit well above the next tier from any other lab. The 1.3-point gap between Opus 4.8 and Gemini 3.1 Pro Preview is narrow enough that future point-release updates from Google could close it.
Access Context
Fable 5 returned to global availability on July 1 after a 19-day suspension under US export controls. It is priced at $10/$50 per million input/output tokens. Mythos 5 — the same underlying weights with safety constraints loosened — remains restricted to Glasswing project partners and approximately 100 government-vetted organizations. The HLE leaderboard reflects Fable 5’s public variant, not the unconstrained Mythos configuration.
GPT-5.6 Sol, which entered restricted government preview on June 26, has not yet appeared on the HLE leaderboard. Its general availability — expected in mid-July — will be the next major data point on whether OpenAI has a credible answer to Fable 5’s expert reasoning scores.