GLM-5.1 Tops Its Open-Weight Class on AA Index at 44 — Frontier Leaders Hold a 23-Point Lead
Z.AI’s GLM-5.1 made headlines on release for topping SWE-Bench Pro ahead of GPT-5.4 and Claude Opus 4.6. The Artificial Analysis Intelligence Index tells a more nuanced story.
The AA Index Placement
On the AA Intelligence Index v4.0, GLM-5.1 (non-reasoning) scores 44 — placing it 28th overall out of 474 tracked models and first in its comparison class of 37 open-weight non-reasoning models. The class median is 21, so GLM-5.1 more than doubles the average for its peer group.
That said, joint frontier leaders GPT-5.4 and Gemini 3.1 Pro Preview both score 57 — a 23-point gap, or roughly 40% above GLM-5.1’s score. Claude Opus 4.6 Adaptive Reasoning sits at 53, GPT-5.3 Codex at 54. GLM-5.1 is not in that tier.
What the AA Index Measures
AA Intelligence Index v4.0 is a 10-evaluation composite: GDPval-AA, τ²-Bench Telecom, Terminal-Bench Hard, SciCode, AA-LCR, AA-Omniscience, IFBench, Humanity’s Last Exam, GPQA Diamond, and CritPt. It spans agentic tasks, scientific reasoning, knowledge, code, and instruction following across 474 models as of April 2026.
SWE-Bench Pro is not part of this composite. Models that are strong coders but weaker on scientific reasoning, agent tasks, or instruction following will score differently across these two benchmarks — and GLM-5.1’s results illustrate the divergence clearly. This matters especially given SWE-Rebench findings from earlier this month showing contamination effects in several models’ SWE-Bench scores.
The Pricing Gap
GLM-5.1 is MIT-licensed open weights — meaning self-hosted inference costs only compute. But via Z.AI’s commercial API, it runs at $1.40 per 1M input tokens and $4.40 per 1M output tokens.
The class average input price for comparable open-weight non-reasoning models is $0.55/M. GLM-5.1’s API pricing puts it at 2.5× that average and in the same tier as proprietary mid-range models — without matching their intelligence scores.
For context, Qwen3.6-Plus from Alibaba offers competitive intelligence at $0.276/M input. The value case for GLM-5.1 via API hinges on use cases where its specific benchmark profile (strong on SWE-type tasks, MIT licensed) justifies the premium.
Model Specs
| Metric | GLM-5.1 |
|---|---|
| AA Intelligence Index | 44 (#28 / 474) |
| Class Rank | #1 / 37 open-weight non-reasoning |
| Input Price | $1.40 / 1M tokens |
| Output Price | $4.40 / 1M tokens |
| Speed | 44 tokens/sec |
| Context | 200K tokens |
| License | MIT |
| Released | April 7, 2026 |
Key Numbers
- AA Intelligence Index score: 44 (frontier leaders: 57)
- Class rank: #1 (class median: 21)
- Overall rank: #28 / 474
- API price: $1.40 input / $4.40 output per 1M tokens (class avg: $0.55 input)
- Cost rank in class: #32 / 37 (among the most expensive)
SWE-Bench Pro first; AA Intelligence Index 28th overall. The model is genuinely capable at the top of its weight class — but benchmark selection shapes the headline considerably.