GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Grok 4.5 Goes Public at $2/$6 Per Million: 62% DeepSWE, 4x Fewer Tokens Than Opus 4.8

SpaceXAI has released Grok 4.5 to the public API at $2/M input and $6/M output tokens. The model targets coding agents and knowledge work, and ships today into Grok Build, Cursor, and the x.ai API. EU availability is not yet confirmed.

Benchmark Position

On DeepSWE 1.0, Grok 4.5 scores 62%, placing third overall behind Claude Fable 5 max and GPT-5.5 xhigh, and ahead of Claude Opus 4.8 max at 55.8%. On Artificial Analysis’s Intelligence Index, it ranks 4th — behind Fable 5, GPT-5.5, and Opus 4.8 — and 4th on GDPval-AA v2.

The model’s primary differentiator is not raw benchmark position. It is token efficiency. On the same Intelligence Index tasks, Grok 4.5 used approximately 60% fewer output tokens than Opus 4.8. Cursor’s own production data shows a similar pattern: Grok 4.5 averaged under 16,000 output tokens per SWE-bench Pro task, roughly 4.2x fewer than Opus 4.8 in the same test.

On practical agentic tasks, the model can build multi-sheet Excel models using web research, design PowerPoint slide content, and write structured prose in Word — suggesting training was weighted toward tool-use patterns rather than benchmark saturation.

The Contamination Disclosure

Cursor has disclosed that an older snapshot of its own codebase was accidentally included in Grok 4.5’s training data, giving the model an unmeasurable advantage on CursorBench. The contaminated data has been removed from future training runs. Cursor’s recommendation: treat its CursorBench numbers for Grok 4.5 as directionally useful but not directly comparable to prior results.

This is the second benchmark transparency event involving Cursor integration data in 2026. The disclosure is a signal that as model developers train on coding-agent data from production platforms, test set contamination will require active monitoring from all parties.

Pricing Context

At $2/M input and $6/M output, Grok 4.5 sits significantly below Opus 4.8 ($5/$25) and Fable 5 ($10/$50). If the token efficiency data holds in production, the effective cost per agent task may be lower still. On the most cost-constrained agentic workflows, Grok 4.5 reaches frontier-tier benchmark scores at roughly one-third the nominal cost of Opus 4.8.

GPT-5.5 at $5/$30 remains the benchmark leader at this class of capability; Grok 4.5’s case is efficiency per dollar, not absolute scores.

Key Numbers

ModelDeepSWE 1.0Input/Output (per 1M)Avg Output Tokens (SWE-bench Pro)
GPT-5.5 xhigh~70%+$5 / $30—
Grok 4.562.0%$2 / $6~16K
Opus 4.8 max55.8%$5 / $25~67K

Free access is available on a limited-time basis inside Grok Build and Cursor, including the x.ai/cli onboarding flow. The 500K-token context window is unchanged from the private beta configuration.