Grok 4.6 Lands in GitHub Copilot — SpaceXAI's 95.6% SWE-Bench Model Reaches Every VS Code Developer
SpaceXAI shipped Grok 4.6 to GitHub Copilot on August 14. Developers using VS Code or any other Copilot surface can now select it from the model picker. The price is unchanged from Grok 4.5: $2 per million input tokens, $6 per million output.
The model underneath is not the same. Grok 4.6 posts 95.6% on SWE-bench Verified under the neutral bash-only harness at vals.ai — ranking fourth globally, behind Claude Opus 5 (97.0%), DeepSeek V4 Pro 0813 (96.4%), and GPT-5.6 Sol (96.2%). Grok 4.5 posted 86.6% on the same benchmark. That 9-point jump moves the model from the second tier into the frontier band where it competes directly with the most capable coding models on the market.
Artificial Analysis puts Grok 4.6’s Intelligence Index at 61, matching GPT-5.6 Sol. Arena added grok-4.6-high to its Code and Text leaderboards on August 12, two days before the Copilot announcement.
The Distribution Math
Copilot’s user base is the relevant variable. SpaceXAI’s own API reaches developers who seek it out. Copilot pushes the model into every default VS Code installation where users have a subscription. The path from “Grok exists” to “Grok is in your IDE” collapses from active adoption to a model picker click.
The pattern follows Grok 4.5’s Copilot integration in July. SpaceXAI is treating the Microsoft distribution as a near-simultaneous launch channel, not an afterthought. Grok 4.5 gained Copilot access roughly three weeks after its public API release. Grok 4.6 launched publicly on August 12 and hit Copilot on August 14 — a 48-hour window.
What 95.6% SWE-Bench Means at $6/M Output
The relevant comparison is cost-per-capability. Claude Fable 5, which posts 95.0% on SWE-bench Verified, costs $50 per million output tokens. Claude Opus 5 at 97.0% costs more still. Grok 4.6 at 95.6% costs $6 per million output — roughly 8x cheaper than Fable 5 at comparable benchmark accuracy.
The caveat: SWE-bench Verified under mini-SWE-agent bash-only conditions is a specific evaluation setup. Production coding agent performance depends on scaffold, tool access, and context handling, where results diverge from the neutral harness. But the benchmark signal is credible: vals.ai runs all models under identical conditions, and the gap between Grok 4.5 and 4.6 on that harness is 9 points.
Arena Signal
SpaceXAI’s Grok 4.6 scores on Arena’s agentic leaderboards remain pending. The model entered Code and Text leaderboards on August 12 but has not yet accumulated enough arena battle results for stable composite agentic rankings. Gemini 3.7 Flash (High) joined Arena’s Agent, Text, and Code leaderboards on August 13 as the other major new entry this week.
The SWE-bench position makes Grok 4.6 the second-cheapest model in the 95%+ tier. GPT-5.6 Sol sits at similar benchmark levels with higher pricing. DeepSeek V4 Pro 0813 at 96.4% runs at substantially lower cost but operates under open-weight licensing constraints for enterprise users. Grok 4.6’s position — frontier-tier SWE-bench, API-hosted, $6/M output, now in Copilot — is a specific combination no other model currently occupies.