GPT-56T 861 —
MUSE-SPK 837 +0.2%
GPT-56SC 790 -4.6%
GLM-5 781 -0.4%
CL-OP55X 780 -5.1%
GROK-46H 780 -5.1%
QWEN-38X 748 -9.2%
GPT-6A 743 -9.4%
KIMI-K3X 742 —
CL-FAB5H 698 -6.1%
CL-OP5H 675 -6.2%
GEM-38FH 672 -0.7%
CL-OP5X 670 -5.5%
CL-OP55H 668 —
CL-OP46H 657 -5.9%
CL-OP47H 648 -6.1%
GPT-56S 618 -0.6%
GEM-37FH 610 -7.2%
GEM-36FH 593 —
CL-OP48H 588 —
CL-OP47 581 -0.2%
GEM-35FH 580 —
GPT-55H 541 -7%
INKL 531 —
GEM-31P 512 -0.2%
CL-OP46 498 +0.4%
GEM-3P 498 -0.2%
CL-OP48 492 +0.4%
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 +0.2%
GPT-56SC 790 -4.6%
GLM-5 781 -0.4%
CL-OP55X 780 -5.1%
GROK-46H 780 -5.1%
QWEN-38X 748 -9.2%
GPT-6A 743 -9.4%
KIMI-K3X 742 —
CL-FAB5H 698 -6.1%
CL-OP5H 675 -6.2%
GEM-38FH 672 -0.7%
CL-OP5X 670 -5.5%
CL-OP55H 668 —
CL-OP46H 657 -5.9%
CL-OP47H 648 -6.1%
GPT-56S 618 -0.6%
GEM-37FH 610 -7.2%
GEM-36FH 593 —
CL-OP48H 588 —
CL-OP47 581 -0.2%
GEM-35FH 580 —
GPT-55H 541 -7%
INKL 531 —
GEM-31P 512 -0.2%
CL-OP46 498 +0.4%
GEM-3P 498 -0.2%
CL-OP48 492 +0.4%
GPT-52 464 —
GPT-55 423 —
← Back to feed

GPT-5.5 Leads Intelligence Index at 60 — and Hallucinates at 86%, the Highest Rate at the Frontier

Artificial Analysis published its full benchmark analysis of GPT-5.5 on April 23, covering all five effort levels. The headline finding: GPT-5.5 (xhigh) takes the Intelligence Index to 60, three points clear of the three-way tie between Anthropic, Google, and OpenAI that held for most of April. The more significant finding is buried in the hallucination data.

The Accuracy/Hallucination Split

GPT-5.5 (xhigh) scores 57% on AA-Omniscience — the highest accuracy AA has ever recorded. It also posts an 86% hallucination rate, meaning the model attempts answers it cannot reliably support at a far higher rate than any other frontier model. For comparison:

  • Claude Opus 4.7 (max): 36% hallucination rate
  • Gemini 3.1 Pro Preview: 50% hallucination rate
  • GPT-5.5 (xhigh): 86% hallucination rate

AA’s framing is precise: GPT-5.5 is more likely than any other model to produce an answer when it does not know one. For knowledge-retrieval applications where confident-but-wrong outputs carry real cost — legal, medical, financial, regulatory — this is a deployment-level constraint, not a benchmark footnote.

The 14-point AA-Omniscience gain from GPT-5.4 (xhigh) was driven primarily by knowledge recall improvements, with only modest movement on the hallucination side. The hallucination problem did not get worse; it was already there in GPT-5.4, and GPT-5.5 did not fix it.

Where It Leads

On the tasks that do not require precision on unknown facts, GPT-5.5 is unambiguously the strongest model on AA’s slate:

  • GDPval-AA Elo: 1785 (xhigh). Leads Claude Opus 4.7 (max) by approximately 30 points and Gemini 3.1 Pro Preview by 470 points. GDPval-AA tests economically valuable real-world tasks.
  • τ²-Bench Telecom: +7 points over GPT-5.4 — the largest gain of any model on this customer-service agent benchmark in the current refresh.
  • Terminal-Bench Hard: leads the field.

The Cost Ladder

Per-token pricing doubled from GPT-5.4 to GPT-5.5 at $5 input / $30 output per 1M tokens. AA’s evaluation found GPT-5.5 uses approximately 40% fewer output tokens than its predecessor, which compresses the net cost increase on the Intelligence Index to approximately 20% more than GPT-5.4.

The more actionable number is the cross-model comparison at the effort level below xhigh:

  • GPT-5.5 (medium): same Intelligence Index score as Claude Opus 4.7 (max). Cost to run AA’s Index: approximately $1,200 — versus $4,800 for Opus 4.7 (max) and $900 for Gemini 3.1 Pro Preview.
  • GPT-5.5 (low): approximates Claude Opus 4.7 (non-reasoning, high). Cost: approximately $500, versus $1,000 for the Opus equivalent.

For use cases where xhigh reasoning depth is not required, the effort ladder makes GPT-5.5 competitive on value despite the higher per-token rate.

The 86% Problem

The hallucination rate creates a hard constraint for production deployments in regulated sectors. A model that recalls facts more accurately than any competitor but attempts answers when it should abstain introduces a class of error that retrieval-augmented generation or grounding can partially mitigate but cannot eliminate from the base model behaviour. Teams building on GPT-5.5 for knowledge-intensive workflows will need factual verification layers that Opus 4.7 users may not require at equivalent scale.

The Intelligence Index leadership at 60 is real. So is the 86%.