GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Grok 4.5 Breaks Into Artificial Analysis Top Tier — From Private Beta to Frontier Benchmarks in 60 Days

Grok 4.5 has a public benchmark score for the first time.

Artificial Analysis’s Intelligence Index now lists Grok 4.5 (high) in fourth position globally, behind Claude Fable 5 (64.9), Claude Opus 4.8 (61.4), and GPT-5.5 (xhigh). When the model entered private beta at SpaceX and Tesla in May, xAI published no benchmark data. The AA Intelligence Index placement is the first independent ranking.

Grok 4.3 scored 53 on the same index in April at 500B parameters. Grok 4.5 runs at roughly 1.5 trillion parameters — triple the scale — after xAI conceded publicly that the Grok 4.3 generation (internal designation V8) was undertrained. The move from 53 into the frontier tier alongside GPT-5.5 and the Claude family marks a significant generation-over-generation jump for xAI’s public benchmark standing.

The “No Benchmarks” Launch

xAI launched Grok 4.5 into private testing without releasing SWE-bench, tau2-bench, or Arena ELO data. The stated cadence was monthly releases with performance claims tied to internal SpaceX and Tesla workloads rather than published evals. At the time, Musk had committed to Grok matching Claude Opus 4.6 performance within that release cycle.

The AA Intelligence Index listing arrives roughly 60 days after that private launch. Grok 4.5’s fourth-place position puts it above Opus 4.6, which is not in the top four.

The “high” qualifier in AA’s listing indicates the reasoning or extended-thinking compute tier of Grok 4.5 — the same distinction AA applies to GPT-5.5 (xhigh) and Claude Opus 4.8 (max). Lower-tier variants of Grok 4.5 may appear on the index separately as evaluations complete.

What’s Still Missing

The AA Intelligence Index is AA’s composite quality metric, not the same as SWE-bench Verified or tau2-bench Telecom performance, which the pipeline uses for agentic composite scoring. Grok 4.5 has not yet published independently verified results on either benchmark.

For context, Grok 4.3 achieved 98.0% on tau2-bench Telecom in April, matching GPT-5.5 on that specific metric. Whether Grok 4.5 extends that result is unknown.

Arena ELO for Grok 4.5 is also absent from the public leaderboards. The model has no Agent Arena entry, and no standard text leaderboard battles have been accumulated. Full competitive positioning will require another 1-2 evaluation cycles.

Structural Note

xAI dissolved as an independent entity following the SpaceX IPO filing. Grok 4.5’s public benchmark debut occurs under the SpaceX AI organisational structure. Whether that changes the publication cadence for safety documentation or system cards is not yet clear.