Grok 4.6 Reaches Intelligence Index 61 — SpaceXAI Matches GPT-5.6 Sol at $2 Per Million
SpaceXAI released Grok 4.6 on August 12, 2026, the company’s first frontier update since Grok 4.5 shipped in late June. The model reaches an Artificial Analysis Intelligence Index score of 61, matching OpenAI’s GPT-5.6 Sol at the frontier tier while holding its price at $2 per million input tokens and $6 per million output tokens.
Where It Lands
The Intelligence Index 61 slots Grok 4.6 into the cluster behind Anthropic’s two leaders:
| Model | AA Intelligence Index | Input / Output per 1M |
|---|---|---|
| Claude Opus 5 (max) | 63 | $5 / $25 |
| Claude Fable 5 (max) | 62 | $15 / $75 |
| Grok 4.6 | 61 | $2 / $6 |
| GPT-5.6 Sol (max) | 61 | $5 / $30 |
| Kimi K3 | ~60 | $3 / $15 |
The 5-point gain over Grok 4.5 (56) came from a longer supplemental training run rather than a parameter increase. The model retains the same 1.5-trillion-parameter V9 foundation but received improved SFT trajectories regenerated by Grok 4.5 across reasoning, coding, and knowledge domains, followed by agentic reinforcement learning on kernel optimisation, web development, and computer-aided design environments.
Agentic Work Is the Story
On GDPval-AA v2, Artificial Analysis’s long-horizon knowledge work evaluation, Grok 4.6 scores 1753 Elo, behind only Claude Opus 5 and within confidence intervals of Fable 5 and Qwen3.8 Max.
On AA-Briefcase, a private long-horizon agentic benchmark, it lands at 1577 Elo — Fable 5 tier — but completes tasks in 53 turns on average against 103 turns for Claude Opus 5. That turn efficiency drives the cost: AA clocks its cost per task at $0.84, the same as Kimi K3.
SpaceXAI’s own evaluation table:
- DeepSWE v1.1: 65.9% (Grok 4.5: 54%; GPT-5.6 Sol Max: 73%)
- Terminal-Bench v2.1: 88.4%
- τ³-Banking: 50.7% (tied with Qwen3.8 Max at the top of this category)
Availability and Pricing
Grok 4.6 is live on the xAI API, Grok Build, Cursor on all plans, OpenRouter, Vercel, and Cloudflare. Context window is 500K tokens, unchanged from Grok 4.5. Cache hits price at $0.50 per million tokens, up from $0.30. SpaceXAI is doubling included usage inside Grok Build and Cursor for the first week.
What It Does Not Claim
No SWE-bench Verified score was published for Grok 4.6. DeepSWE v1.1 (65.9%) uses a different task distribution and harness from SWE-bench Verified, making direct comparison to Fable 5’s 95% or Opus 5’s 97% unreliable. The benchmark figures in the SpaceXAI launch table are self-reported; no independent evaluation has confirmed them.
The strategic position is clear regardless: frontier-tier general capability at a price point 60% below the models it now ties.