GLM-52 883 -1.9%
GROK-45 870 -1.8%
GPT-56T 861
GROK-46H 856 -2.8%
DSK-V4PH 856 -2.8%
DSK-V4FH 856 -2.8%
CL-OP48H 856 -2.8%
CL-OP5X 842 -8.3%
QWEN-38X 840
MUSE-SPK 829 -0.6%
GPT-6A 820
CL-OP5H 807 -1%
GPT-56SC 787
GLM-5 781 -8.4%
CL-FAB5H 777 -3.1%
CL-OP46H 772 -3.1%
CL-OP47H 770 -3.1%
GEM-37FH 748 -15.1%
KIMI-K3X 742 -9.5%
GPT-56S 710
GEM-38FH 709 -1.1%
CL-OP47 659 -0.9%
CL-OP48 617 -0.2%
GPT-55H 612
INKL 531
GEM-31P 519
GEM-3P 501
CL-OP46 499 -0.2%
GLM-52 883 -1.9%
GROK-45 870 -1.8%
GPT-56T 861
GROK-46H 856 -2.8%
DSK-V4PH 856 -2.8%
DSK-V4FH 856 -2.8%
CL-OP48H 856 -2.8%
CL-OP5X 842 -8.3%
QWEN-38X 840
MUSE-SPK 829 -0.6%
GPT-6A 820
CL-OP5H 807 -1%
GPT-56SC 787
GLM-5 781 -8.4%
CL-FAB5H 777 -3.1%
CL-OP46H 772 -3.1%
CL-OP47H 770 -3.1%
GEM-37FH 748 -15.1%
KIMI-K3X 742 -9.5%
GPT-56S 710
GEM-38FH 709 -1.1%
CL-OP47 659 -0.9%
CL-OP48 617 -0.2%
GPT-55H 612
INKL 531
GEM-31P 519
GEM-3P 501
CL-OP46 499 -0.2%
← Back to feed

xAI Confirms Grok V9 at 1.5T Parameters — Foundation Model Triples in Scale as Musk Admits V8 Was Undertrained

Elon Musk posted directly on X confirming xAI’s next foundation model: version 9, at 1.5 trillion parameters, optimized for Blackwell GPUs. The current public-facing Grok 4.2 runs on foundation model v8 at 0.5 trillion parameters — trained on Hopper hardware, and per Musk, carrying “significant shortfalls in training data quality, comprehensiveness and proportionality.”

The jump from v8 to v9 is a 3x parameter increase. Musk’s framing was blunt: “The difference between Grok foundation model 8 and 9 is gigantic.”

What Musk Said, Exactly

The version numbering confusion Musk addressed publicly is meaningful. The internal v-number tracks foundation model generations; the public x.y version tracks product releases. Grok 4.2 = foundation v8. Whatever ships next based on v9 will likely arrive as Grok 4.4 or 4.5 before the Grok 5 family targeting 6–10 trillion parameters.

Key admissions in the post:

  • V8 is “only 0.5T in size”
  • V8 was “trained on Hoppers” — the older GPU generation
  • V8 has “shortfalls in training data quality, comprehensiveness and proportionality”
  • V9 is “substantially better in every way: data curation, training recipe, size”
  • V9 is “optimized to run on Blackwells”

This is a rare instance of a lab founder explicitly characterising a shipping model as deficient on its own terms.

Context: Grok Is Losing Ground

The Wall Street Journal reported separately this week that Grok is losing ground in the AI race. Usage data and Arena standings confirm the gap. Grok 4.3 — the current flagship, scored at 53 on the Artificial Analysis Intelligence Index — sits well behind GPT-5.5 (60), Claude Opus 4.7 (58), and Gemini 3.1 Pro Preview (57). On Terminal-Bench 2.0, Grok doesn’t appear in the top 15.

The competitive deficit Musk is racing against is real. xAI’s compute access is also real: Anthropic’s deal to take all 220,000 GPUs at SpaceX’s Colossus 1 means xAI must build its own next-generation cluster, which it is doing. Colossus 2 is reportedly targeting multiple gigawatts.

The Roadmap

Based on Musk’s post and independent reporting, the xAI foundation model ladder currently looks like this:

Foundation VersionParametersPublic ProductHardware
v80.5TGrok 4.2 (current)Hopper
v91.5TGrok 4.4–4.5 (next)Blackwell
Grok 56–10TTBDColossus 2

The 10 trillion parameter Grok 5 run is reportedly two months of pre-training on xAI’s Colossus cluster. Whether that timeline holds is an open question; Musk has missed model release deadlines across the Grok family.

What This Changes

xAI’s public acknowledgment of v8’s deficiencies shifts how Grok 4.x performance should be read. Benchmark scores for the current product are not representative of what the underlying architecture can achieve — they reflect a training execution that Musk himself has flagged as flawed. V9 with Blackwell-native training and corrected data may land substantially higher on intelligence and coding benchmarks when it ships, potentially in Q3 2026. Whether it reaches parity with Claude Opus 4.7 or GPT-5.5 by then depends on how much of the gap is scale versus architecture.

At 1.5T parameters, v9 is entering territory where raw compute starts to matter less than data quality and RL training. That’s exactly what Musk says they’re fixing.