GPT-5.5 Leads Intelligence Index at 60 — and Hallucinates at 86%, the Highest Rate at the Frontier
Artificial Analysis published its full benchmark analysis of GPT-5.5 on April 23, covering all five effort levels. The headline finding: GPT-5.5 (xhigh) takes the Intelligence Index to 60, three points clear of the three-way tie between Anthropic, Google, and OpenAI that held for most of April. The more significant finding is buried in the hallucination data.
The Accuracy/Hallucination Split
GPT-5.5 (xhigh) scores 57% on AA-Omniscience — the highest accuracy AA has ever recorded. It also posts an 86% hallucination rate, meaning the model attempts answers it cannot reliably support at a far higher rate than any other frontier model. For comparison:
- Claude Opus 4.7 (max): 36% hallucination rate
- Gemini 3.1 Pro Preview: 50% hallucination rate
- GPT-5.5 (xhigh): 86% hallucination rate
AA’s framing is precise: GPT-5.5 is more likely than any other model to produce an answer when it does not know one. For knowledge-retrieval applications where confident-but-wrong outputs carry real cost — legal, medical, financial, regulatory — this is a deployment-level constraint, not a benchmark footnote.
The 14-point AA-Omniscience gain from GPT-5.4 (xhigh) was driven primarily by knowledge recall improvements, with only modest movement on the hallucination side. The hallucination problem did not get worse; it was already there in GPT-5.4, and GPT-5.5 did not fix it.
Where It Leads
On the tasks that do not require precision on unknown facts, GPT-5.5 is unambiguously the strongest model on AA’s slate:
- GDPval-AA Elo: 1785 (xhigh). Leads Claude Opus 4.7 (max) by approximately 30 points and Gemini 3.1 Pro Preview by 470 points. GDPval-AA tests economically valuable real-world tasks.
- τ²-Bench Telecom: +7 points over GPT-5.4 — the largest gain of any model on this customer-service agent benchmark in the current refresh.
- Terminal-Bench Hard: leads the field.
The Cost Ladder
Per-token pricing doubled from GPT-5.4 to GPT-5.5 at $5 input / $30 output per 1M tokens. AA’s evaluation found GPT-5.5 uses approximately 40% fewer output tokens than its predecessor, which compresses the net cost increase on the Intelligence Index to approximately 20% more than GPT-5.4.
The more actionable number is the cross-model comparison at the effort level below xhigh:
- GPT-5.5 (medium): same Intelligence Index score as Claude Opus 4.7 (max). Cost to run AA’s Index: approximately $1,200 — versus $4,800 for Opus 4.7 (max) and $900 for Gemini 3.1 Pro Preview.
- GPT-5.5 (low): approximates Claude Opus 4.7 (non-reasoning, high). Cost: approximately $500, versus $1,000 for the Opus equivalent.
For use cases where xhigh reasoning depth is not required, the effort ladder makes GPT-5.5 competitive on value despite the higher per-token rate.
The 86% Problem
The hallucination rate creates a hard constraint for production deployments in regulated sectors. A model that recalls facts more accurately than any competitor but attempts answers when it should abstain introduces a class of error that retrieval-augmented generation or grounding can partially mitigate but cannot eliminate from the base model behaviour. Teams building on GPT-5.5 for knowledge-intensive workflows will need factual verification layers that Opus 4.7 users may not require at equivalent scale.
The Intelligence Index leadership at 60 is real. So is the 86%.