GPT-56T 861 —
MUSE-SPK 837 +0.2%
GPT-56SC 790 -4.6%
GLM-5 781 -0.4%
CL-OP55X 780 -5.1%
GROK-46H 780 -5.1%
QWEN-38X 748 -9.2%
GPT-6A 743 -9.4%
KIMI-K3X 742 —
CL-FAB5H 698 -6.1%
CL-OP5H 675 -6.2%
GEM-38FH 672 -0.7%
CL-OP5X 670 -5.5%
CL-OP55H 668 —
CL-OP46H 657 -5.9%
CL-OP47H 648 -6.1%
GPT-56S 618 -0.6%
GEM-37FH 610 -7.2%
GEM-36FH 593 —
CL-OP48H 588 —
CL-OP47 581 -0.2%
GEM-35FH 580 —
GPT-55H 541 -7%
INKL 531 —
GEM-31P 512 -0.2%
CL-OP46 498 +0.4%
GEM-3P 498 -0.2%
CL-OP48 492 +0.4%
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 +0.2%
GPT-56SC 790 -4.6%
GLM-5 781 -0.4%
CL-OP55X 780 -5.1%
GROK-46H 780 -5.1%
QWEN-38X 748 -9.2%
GPT-6A 743 -9.4%
KIMI-K3X 742 —
CL-FAB5H 698 -6.1%
CL-OP5H 675 -6.2%
GEM-38FH 672 -0.7%
CL-OP5X 670 -5.5%
CL-OP55H 668 —
CL-OP46H 657 -5.9%
CL-OP47H 648 -6.1%
GPT-56S 618 -0.6%
GEM-37FH 610 -7.2%
GEM-36FH 593 —
CL-OP48H 588 —
CL-OP47 581 -0.2%
GEM-35FH 580 —
GPT-55H 541 -7%
INKL 531 —
GEM-31P 512 -0.2%
CL-OP46 498 +0.4%
GEM-3P 498 -0.2%
CL-OP48 492 +0.4%
GPT-52 464 —
GPT-55 423 —
← Back to feed

Grok 4.3 Official Release Scores 53 on AA Intelligence Index at $2.50/M Output

Grok 4.3 received its full Artificial Analysis benchmark entry on April 30, six weeks after the beta launched to SuperGrok Heavy subscribers. The model posts an AA Intelligence Index of 53, placing it 10th out of 154 evaluated models.

Key Numbers

  • AA Intelligence Index: 53 (#10/154)
  • Input cost: $1.25/M tokens
  • Output cost: $2.50/M tokens
  • Context window: 1M tokens
  • Modalities: Text and image input, text output
  • Output verbosity: 88M tokens generated during AA evaluation — 2.5× the 35M median for comparable models

Where It Sits on the Frontier

The 53 score clusters Grok 4.3 with Kimi K2.6 (54) and Xiaomi MiMo-V2.5-Pro (54), 7 points below GPT-5.5 (60) and Claude Opus 4.7 at the current ceiling. It sits above Qwen3.6 Max Preview (52) and comfortably ahead of Gemini 2.5 Pro class models.

The pricing spread tells a larger story. At $2.50/M output, Grok 4.3 undercuts Claude Opus 4.7 ($150/M output) by 98%, GPT-5.5-High by a comparable margin, and sits cheaper than Kimi K2.6 for most workload profiles. For applications that can tolerate a 7-point Intelligence Index gap, the economics shift sharply.

The Verbosity Problem

The 88M-token output volume is a structural concern. During AA evaluation, Grok 4.3 generated 2.5× more tokens than the median for reasoning models in its price tier. At $2.50/M output, a verbose-by-default model’s effective per-query cost is meaningfully higher than the headline rate suggests — and for token-counted APIs, verbosity is a cost multiplier the pricing sheet does not show.

xAI has not published a system prompt or configuration that reduces output length.

Context Window Revision

The beta launched in April with reports of a 2M token context window. The production model on Artificial Analysis shows 1M tokens — the same as GPT-4.1, below Kimi K2.6 (131K) or Claude Opus 4.7 (200K). Whether the 2M figure was a beta capability that was removed, or an error in early reporting, has not been clarified by xAI.

Benchmark Composition

Artificial Analysis Intelligence Index v4.0 incorporates 10 evaluations: GDPval-AA, τ²-Bench Telecom, Terminal-Bench Hard, SciCode, AA-LCR, AA-Omniscience, IFBench, Humanity’s Last Exam, GPQA Diamond, and CritPt. Grok 4.3 has not published separate SWE-bench Verified or independent tau2-bench results. The AA composite is currently the only third-party benchmark on record for this model.

Total evaluation cost at AA: $395.17 — higher than average, consistent with the token verbosity.

SpaceXAI Cadence

xAI confirmed a two-week model release cadence from its Colossus facility. Grok 4.3’s full release is the formal close of the April cycle. A successor is expected before mid-May.