GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
← Back to feed

GLM-5.1 Tops Its Open-Weight Class on AA Index at 44 — Frontier Leaders Hold a 23-Point Lead

Z.AI’s GLM-5.1 made headlines on release for topping SWE-Bench Pro ahead of GPT-5.4 and Claude Opus 4.6. The Artificial Analysis Intelligence Index tells a more nuanced story.

The AA Index Placement

On the AA Intelligence Index v4.0, GLM-5.1 (non-reasoning) scores 44 — placing it 28th overall out of 474 tracked models and first in its comparison class of 37 open-weight non-reasoning models. The class median is 21, so GLM-5.1 more than doubles the average for its peer group.

That said, joint frontier leaders GPT-5.4 and Gemini 3.1 Pro Preview both score 57 — a 23-point gap, or roughly 40% above GLM-5.1’s score. Claude Opus 4.6 Adaptive Reasoning sits at 53, GPT-5.3 Codex at 54. GLM-5.1 is not in that tier.

What the AA Index Measures

AA Intelligence Index v4.0 is a 10-evaluation composite: GDPval-AA, τ²-Bench Telecom, Terminal-Bench Hard, SciCode, AA-LCR, AA-Omniscience, IFBench, Humanity’s Last Exam, GPQA Diamond, and CritPt. It spans agentic tasks, scientific reasoning, knowledge, code, and instruction following across 474 models as of April 2026.

SWE-Bench Pro is not part of this composite. Models that are strong coders but weaker on scientific reasoning, agent tasks, or instruction following will score differently across these two benchmarks — and GLM-5.1’s results illustrate the divergence clearly. This matters especially given SWE-Rebench findings from earlier this month showing contamination effects in several models’ SWE-Bench scores.

The Pricing Gap

GLM-5.1 is MIT-licensed open weights — meaning self-hosted inference costs only compute. But via Z.AI’s commercial API, it runs at $1.40 per 1M input tokens and $4.40 per 1M output tokens.

The class average input price for comparable open-weight non-reasoning models is $0.55/M. GLM-5.1’s API pricing puts it at 2.5× that average and in the same tier as proprietary mid-range models — without matching their intelligence scores.

For context, Qwen3.6-Plus from Alibaba offers competitive intelligence at $0.276/M input. The value case for GLM-5.1 via API hinges on use cases where its specific benchmark profile (strong on SWE-type tasks, MIT licensed) justifies the premium.

Model Specs

MetricGLM-5.1
AA Intelligence Index44 (#28 / 474)
Class Rank#1 / 37 open-weight non-reasoning
Input Price$1.40 / 1M tokens
Output Price$4.40 / 1M tokens
Speed44 tokens/sec
Context200K tokens
LicenseMIT
ReleasedApril 7, 2026

Key Numbers

  • AA Intelligence Index score: 44 (frontier leaders: 57)
  • Class rank: #1 (class median: 21)
  • Overall rank: #28 / 474
  • API price: $1.40 input / $4.40 output per 1M tokens (class avg: $0.55 input)
  • Cost rank in class: #32 / 37 (among the most expensive)

SWE-Bench Pro first; AA Intelligence Index 28th overall. The model is genuinely capable at the top of its weight class — but benchmark selection shapes the headline considerably.