GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

GLM-5.2 Leads Open-Weights Intelligence Index at 51 — Matches GPT-5.5 on Real-World Agent Tasks

Z.ai’s GLM-5.2 now leads every open-weight model on Artificial Analysis’s Intelligence Index v4.1, scoring 51 against the next-closest rivals at 44. More striking: on GDPval-AA v2, the real-world agent benchmark that AA treats as its primary metric, GLM-5.2 scores 1524 — effectively level with GPT-5.5 xhigh at 1514.

That comparison invites scrutiny, but the benchmark is designed for it. GDPval-AA v2 baselines Elo to human performance at 1000, uses a rotating panel of frontier-model judges, and raises the turn limit from 100 to 250 for longer-horizon trajectories. GLM-5.2’s 1524 places ahead of MiniMax-M3 (1418) and DeepSeek V4 Pro max (1328).

Intelligence Index v4.1 Rankings

ModelIntelligence IndexGDPval-AA v2
GLM-5.2511524
MiniMax-M3441418
DeepSeek V4 Pro (max)441328
Kimi K2.643—

The index incorporates nine evaluations: GDPval-AA v2, tau3-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR. GLM-5.2’s gains over its predecessor are spread across most of them.

Where GLM-5.2 Gained on GLM-5.1

  • Terminal-Bench v2.1: 78% (+16 points)
  • tau3-Banking: 27% (+15 points)
  • HLE (Humanity’s Last Exam): 40% (+12 points)
  • AA-LCR (long-context reasoning): 71% (+9 points)
  • SciCode: 50% (+7 points)
  • GPQA Diamond: 89% (+3 points)

The scientific reasoning gains are notable — HLE and CritPt both improved sharply. Terminal-Bench performance at 78% also puts GLM-5.2 ahead of most proprietary models outside the Claude Fable 5 / GPT-5.5 tier.

Architecture and Pricing

GLM-5.2 uses the same architecture as GLM-5.1: 744B total parameters, 40B active. The context window extends from 200K to 1M tokens. Pricing is unchanged at $1.40 per 1M input tokens and $4.40 per 1M output tokens, with cached input at $0.26 per 1M. MIT license.

What changed is verbosity. GLM-5.2 uses 43K output tokens per Intelligence Index task on average, up from 26K for GLM-5.1, and higher than MiniMax-M3 (24K), Kimi K2.6 (35K), and DeepSeek V4 Pro max (37K). That reasoning budget buys the benchmark gains, but puts GLM-5.2 toward the cost-per-task ceiling among open-weight models. At roughly $0.46 per task, it costs more than Kimi K2.6 ($0.31) and considerably more than DeepSeek V4 Pro max ($0.05).

On the AA cost vs. intelligence Pareto frontier, GLM-5.2 sits at the edge of the most attractive quadrant — it’s the cheapest model at its intelligence level among proprietary-tier performers, but not cheap compared to the wider open-weights field.

Availability

Available on the Z.ai first-party API and across DeepInfra, Novita, Nebius, Parasail, Siliconflow, GMI Cloud, Baseten, and Fireworks. GLM-5.2 entered the Arena Text and Code leaderboards on June 16; the Agent Arena ranking is pending sufficient vote accumulation.

The open-weights intelligence gap to proprietary frontiers has been 10-20 points for most of 2026. GLM-5.2 narrows it to 13 against GPT-5.5 high (64 on Intelligence Index) and to 14 against Claude Opus 4.8. For teams that can accept MIT-licensed weights and the associated computational cost, that margin is now arguably workable.