GLM-5.2 Leads Open-Weights Intelligence Index at 51 — Matches GPT-5.5 on Real-World Agent Tasks
Z.ai’s GLM-5.2 now leads every open-weight model on Artificial Analysis’s Intelligence Index v4.1, scoring 51 against the next-closest rivals at 44. More striking: on GDPval-AA v2, the real-world agent benchmark that AA treats as its primary metric, GLM-5.2 scores 1524 — effectively level with GPT-5.5 xhigh at 1514.
That comparison invites scrutiny, but the benchmark is designed for it. GDPval-AA v2 baselines Elo to human performance at 1000, uses a rotating panel of frontier-model judges, and raises the turn limit from 100 to 250 for longer-horizon trajectories. GLM-5.2’s 1524 places ahead of MiniMax-M3 (1418) and DeepSeek V4 Pro max (1328).
Intelligence Index v4.1 Rankings
| Model | Intelligence Index | GDPval-AA v2 |
|---|---|---|
| GLM-5.2 | 51 | 1524 |
| MiniMax-M3 | 44 | 1418 |
| DeepSeek V4 Pro (max) | 44 | 1328 |
| Kimi K2.6 | 43 | — |
The index incorporates nine evaluations: GDPval-AA v2, tau3-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR. GLM-5.2’s gains over its predecessor are spread across most of them.
Where GLM-5.2 Gained on GLM-5.1
- Terminal-Bench v2.1: 78% (+16 points)
- tau3-Banking: 27% (+15 points)
- HLE (Humanity’s Last Exam): 40% (+12 points)
- AA-LCR (long-context reasoning): 71% (+9 points)
- SciCode: 50% (+7 points)
- GPQA Diamond: 89% (+3 points)
The scientific reasoning gains are notable — HLE and CritPt both improved sharply. Terminal-Bench performance at 78% also puts GLM-5.2 ahead of most proprietary models outside the Claude Fable 5 / GPT-5.5 tier.
Architecture and Pricing
GLM-5.2 uses the same architecture as GLM-5.1: 744B total parameters, 40B active. The context window extends from 200K to 1M tokens. Pricing is unchanged at $1.40 per 1M input tokens and $4.40 per 1M output tokens, with cached input at $0.26 per 1M. MIT license.
What changed is verbosity. GLM-5.2 uses 43K output tokens per Intelligence Index task on average, up from 26K for GLM-5.1, and higher than MiniMax-M3 (24K), Kimi K2.6 (35K), and DeepSeek V4 Pro max (37K). That reasoning budget buys the benchmark gains, but puts GLM-5.2 toward the cost-per-task ceiling among open-weight models. At roughly $0.46 per task, it costs more than Kimi K2.6 ($0.31) and considerably more than DeepSeek V4 Pro max ($0.05).
On the AA cost vs. intelligence Pareto frontier, GLM-5.2 sits at the edge of the most attractive quadrant — it’s the cheapest model at its intelligence level among proprietary-tier performers, but not cheap compared to the wider open-weights field.
Availability
Available on the Z.ai first-party API and across DeepInfra, Novita, Nebius, Parasail, Siliconflow, GMI Cloud, Baseten, and Fireworks. GLM-5.2 entered the Arena Text and Code leaderboards on June 16; the Agent Arena ranking is pending sufficient vote accumulation.
The open-weights intelligence gap to proprietary frontiers has been 10-20 points for most of 2026. GLM-5.2 narrows it to 13 against GPT-5.5 high (64 on Intelligence Index) and to 14 against Claude Opus 4.8. For teams that can accept MIT-licensed weights and the associated computational cost, that margin is now arguably workable.