GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
← Back to feed

Musk Sets June Deadline for Grok to Match Claude Opus 4.6 — After Admitting xAI Was "Not Built Right"

Elon Musk publicly set a competitive clock on April 11: Grok will approach Claude Opus 4.6 performance by May 2026, match or exceed it by June. The claim — posted directly on X — comes as xAI is mid-rebuild after Musk acknowledged in March that the company “was not built right first time around.”

The context makes the timeline aggressive. Grok 4.1, the current production model, posts a 15.9% score on ARC-AGI V2. Gemini 3.1 Pro sits at 77.1% on the same benchmark. Claude Opus 4.6 scores 41.2%. The gap on abstract reasoning is substantial, and it is the kind of gap that architectural patches do not close without a generational model release.

What xAI Is Actually Building

The June target almost certainly refers to Grok 5, currently in training at roughly 6 trillion parameters — compared to the 1.7 trillion of Grok 4. Musk has placed a 10% probability on Grok 5 achieving AGI-level capability, a figure he has been consistent about. If the training timeline holds, Grok 5 would be the first xAI model built on the new architecture.

xAI ended 2025 with over one million NVIDIA H100 GPU equivalents across its Colossus I and II supercomputers. The $20 billion Series E gives it the compute runway, but the architectural rebuild — while simultaneously shipping products and integrating with SpaceX post-merger — is a parallel execution challenge.

Current Grok Performance vs the Field

BenchmarkGrok 4.1Claude Opus 4.6Gemini 3.1 ProGPT-5.4 High
GPQA Diamond~87%91.3%94.3%92.4%
ARC-AGI V215.9%37.6%77.1%52.9%
HLE (no tools)~25%41.2%44.4%34.5%

On math, Grok 4 Heavy was the first model past 50% on HLE at launch in July 2025. By April 2026, no competitor has crossed that threshold, but Gemini 3.1 Pro (44.4%) and Claude Opus 4.6 (41.2%) have narrowed the margin substantially.

The IPO Forcing Function

xAI has a June 2026 IPO target for the combined SpaceX-xAI entity, valued at approximately $250 billion as part of a $1.25 trillion combined listing. Showing a competitive Grok 5 before that window would directly support the public market narrative. A June model release and a June IPO sitting on the same timeline is not coincidence.

What “Match and Maybe Exceed” Actually Means

Musk’s framing — “close to Opus 4.6” by May, “match and maybe exceed” by June — is careful. Claude Opus 4.6 holds an 80.8% SWE-bench Verified score and strong tau-bench retail performance (91.9%). Those agentic benchmarks require sustained, multi-step software engineering under realistic conditions. Matching them with a model fresh off an architectural rebuild, within two months, is the kind of target that either redefines the AI generation or becomes a post-IPO clarification moment.

The architectural rebuild is real. The compute is real. The June deadline is Musk’s to keep.