First-Ever Three-Way Tie at Intelligence Index 57: All Three Frontier Labs Now Equal
Artificial Analysis has announced the first-ever equal-first finish on its Intelligence Index: Claude Opus 4.7 (57.3), Gemini 3.1 Pro Preview (57.2), and GPT-5.4 (56.8) all fall within the benchmark’s 95% confidence interval of ±1 point. AA explicitly recommends treating these as a tie.
The index aggregates 10 evaluations: GDPval-AA, τ²-Bench Telecom, Terminal-Bench Hard, SciCode, AA-LCR, AA-Omniscience, IFBench, HLE, GPQA Diamond, and CritPt.
The Differentiation Underneath the Tie
The equal composite score masks clean lab-level specialisation that now makes model selection a question of workload, not headline score.
Anthropic leads real-world agentic work. Claude Opus 4.7 tops GDPval-AA, AA’s primary benchmark for general agentic performance across 44 occupations and 9 industries. It also leads the Omniscience Index at #2 (behind Gemini 3.1 Pro) with measurably lower hallucination rates than Opus 4.6.
Google leads scientific and knowledge reasoning. Gemini 3.1 Pro Preview tops HLE, GPQA Diamond, SciCode, IFBench, and AA-Omniscience — the full sweep of knowledge-intensive evaluations. The highest-knowledge workloads consistently route to Google.
OpenAI leads long-horizon coding and scientific reasoning. GPT-5.4 tops Terminal-Bench Hard, CritPt, and AA-LCR — the longest and most complex coding and reasoning chains. For agentic coding over extended workflows, OpenAI retains a measurable edge.
What This Means
Until April 2026, there was always a clear #1 on the Intelligence Index. Anthropic or OpenAI held it, with the other close behind. The equal-first result signals something structurally different: the three frontier labs have independently converged on the same capability ceiling, reached from different architectural and training directions.
The practical consequence is that single-metric model selection is no longer valid at the frontier. A team using Opus 4.7 for enterprise agentic work and Gemini 3.1 Pro for scientific Q&A is making a better decision than one that picks a single model based on a composite rank that no longer differentiates.
What Comes Next
Anthropic’s Claude Mythos Preview — still in limited access, with a reported 93.9% SWE-bench Verified score — sits above all three in coding capability but is not yet generally available. Whether Mythos shifts the AA index when it ships broadly is the question that will end the tie.
The ceiling for the Intelligence Index, currently at 57 across three models, has never been held by this many labs simultaneously.