Grok 4.3 Official Release Scores 53 on AA Intelligence Index at $2.50/M Output
Grok 4.3 received its full Artificial Analysis benchmark entry on April 30, six weeks after the beta launched to SuperGrok Heavy subscribers. The model posts an AA Intelligence Index of 53, placing it 10th out of 154 evaluated models.
Key Numbers
- AA Intelligence Index: 53 (#10/154)
- Input cost: $1.25/M tokens
- Output cost: $2.50/M tokens
- Context window: 1M tokens
- Modalities: Text and image input, text output
- Output verbosity: 88M tokens generated during AA evaluation — 2.5× the 35M median for comparable models
Where It Sits on the Frontier
The 53 score clusters Grok 4.3 with Kimi K2.6 (54) and Xiaomi MiMo-V2.5-Pro (54), 7 points below GPT-5.5 (60) and Claude Opus 4.7 at the current ceiling. It sits above Qwen3.6 Max Preview (52) and comfortably ahead of Gemini 2.5 Pro class models.
The pricing spread tells a larger story. At $2.50/M output, Grok 4.3 undercuts Claude Opus 4.7 ($150/M output) by 98%, GPT-5.5-High by a comparable margin, and sits cheaper than Kimi K2.6 for most workload profiles. For applications that can tolerate a 7-point Intelligence Index gap, the economics shift sharply.
The Verbosity Problem
The 88M-token output volume is a structural concern. During AA evaluation, Grok 4.3 generated 2.5× more tokens than the median for reasoning models in its price tier. At $2.50/M output, a verbose-by-default model’s effective per-query cost is meaningfully higher than the headline rate suggests — and for token-counted APIs, verbosity is a cost multiplier the pricing sheet does not show.
xAI has not published a system prompt or configuration that reduces output length.
Context Window Revision
The beta launched in April with reports of a 2M token context window. The production model on Artificial Analysis shows 1M tokens — the same as GPT-4.1, below Kimi K2.6 (131K) or Claude Opus 4.7 (200K). Whether the 2M figure was a beta capability that was removed, or an error in early reporting, has not been clarified by xAI.
Benchmark Composition
Artificial Analysis Intelligence Index v4.0 incorporates 10 evaluations: GDPval-AA, τ²-Bench Telecom, Terminal-Bench Hard, SciCode, AA-LCR, AA-Omniscience, IFBench, Humanity’s Last Exam, GPQA Diamond, and CritPt. Grok 4.3 has not published separate SWE-bench Verified or independent tau2-bench results. The AA composite is currently the only third-party benchmark on record for this model.
Total evaluation cost at AA: $395.17 — higher than average, consistent with the token verbosity.
SpaceXAI Cadence
xAI confirmed a two-week model release cadence from its Colossus facility. Grok 4.3’s full release is the formal close of the April cycle. A successor is expected before mid-May.