GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
← Back to feed

First-Ever Three-Way Tie at Intelligence Index 57: All Three Frontier Labs Now Equal

Artificial Analysis has announced the first-ever equal-first finish on its Intelligence Index: Claude Opus 4.7 (57.3), Gemini 3.1 Pro Preview (57.2), and GPT-5.4 (56.8) all fall within the benchmark’s 95% confidence interval of ±1 point. AA explicitly recommends treating these as a tie.

The index aggregates 10 evaluations: GDPval-AA, τ²-Bench Telecom, Terminal-Bench Hard, SciCode, AA-LCR, AA-Omniscience, IFBench, HLE, GPQA Diamond, and CritPt.

The Differentiation Underneath the Tie

The equal composite score masks clean lab-level specialisation that now makes model selection a question of workload, not headline score.

Anthropic leads real-world agentic work. Claude Opus 4.7 tops GDPval-AA, AA’s primary benchmark for general agentic performance across 44 occupations and 9 industries. It also leads the Omniscience Index at #2 (behind Gemini 3.1 Pro) with measurably lower hallucination rates than Opus 4.6.

Google leads scientific and knowledge reasoning. Gemini 3.1 Pro Preview tops HLE, GPQA Diamond, SciCode, IFBench, and AA-Omniscience — the full sweep of knowledge-intensive evaluations. The highest-knowledge workloads consistently route to Google.

OpenAI leads long-horizon coding and scientific reasoning. GPT-5.4 tops Terminal-Bench Hard, CritPt, and AA-LCR — the longest and most complex coding and reasoning chains. For agentic coding over extended workflows, OpenAI retains a measurable edge.

What This Means

Until April 2026, there was always a clear #1 on the Intelligence Index. Anthropic or OpenAI held it, with the other close behind. The equal-first result signals something structurally different: the three frontier labs have independently converged on the same capability ceiling, reached from different architectural and training directions.

The practical consequence is that single-metric model selection is no longer valid at the frontier. A team using Opus 4.7 for enterprise agentic work and Gemini 3.1 Pro for scientific Q&A is making a better decision than one that picks a single model based on a composite rank that no longer differentiates.

What Comes Next

Anthropic’s Claude Mythos Preview — still in limited access, with a reported 93.9% SWE-bench Verified score — sits above all three in coding capability but is not yet generally available. Whether Mythos shifts the AA index when it ships broadly is the question that will end the tie.

The ceiling for the Intelligence Index, currently at 57 across three models, has never been held by this many labs simultaneously.