GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

GLM 5.2 Debuts at #10 on Agent Arena — Z.ai's Open-Weight Model Cracks the Proprietary Tier

Z.ai’s GLM 5.2 (Max) was added to the Agent Arena leaderboard on June 16 and debuted at rank 10, placing it inside a tier previously occupied entirely by closed proprietary models from Anthropic, OpenAI, and Google.

Its composite score of 4.37% (±2.48%) puts it above Claude Opus 4.8 standard (rank 11, 3.60%) and Claude Sonnet 4.6 (rank 12, 3.22%). The comparison is imperfect — GLM 5.2 has only 9,625 votes, roughly one-third the vote count of the established models, which explains the wider uncertainty margin. Position will stabilise as more battles accumulate.

The context: GLM 5.1 sits at rank 13 with 2.66%. GLM 5.2’s debut at 10 represents a meaningful generational jump for the Z.ai model family.

Category Breakdown

The composite conceals some category-level variance worth noting:

CategoryGLM 5.2 (Max)Claude Opus 4.8 Thinking
Coding9.43%10.75%
Web14.88%14.36%
Computer Use6.00%10.48%
Research1.69%9.08%
Agent Safety1.86%0.56%

Web tasks are GLM 5.2’s strongest category (14.88%), where it edges Claude Opus 4.8 Thinking (14.36%). Research tasks (1.69%) represent its weakest showing — well below the Anthropic and OpenAI frontier.

MiniMax M3 Also Enters

MiniMax M3 was added to the Agent Arena on the same day and debuted at rank 19 with a 2.79% composite (±1.70%), 10,258 votes. It places below GLM 5.2 and above Qwen 3.6 Plus (rank 20, 4.24% with 32,121 votes). Given MiniMax M3’s higher vote count, Qwen 3.6 Plus’s lower position is likely more reliable at this sample size.

Open-Weight Position

With GLM 5.2 at rank 10, Z.ai holds two positions in the top 13 (GLM 5.1 at rank 13). Both models carry MIT licenses and route through SiliconFlow. The Agent Arena top 10 is now:

  1. Claude Fable 5 (High) — 14.17%
  2. Claude Opus 4.8 (Thinking) — 9.04%
  3. GPT-5.5 (xHigh) — 8.27% 4–5. Claude Opus 4.7 / 4.7 Thinking — ~8.1%
  4. GPT-5.5 (High) — 7.78%
  5. GPT-5.5 — 6.73%
  6. Claude Opus 4.6 — 6.73%
  7. GPT-5.4 (High) — 6.54%
  8. GLM 5.2 (Max) — 4.37% (MIT, open-weight)

A 2.17 point gap separates GPT-5.4 High (#9) from GLM 5.2 (#10). That gap will narrow or widen as vote counts equalise. The ranking is live and continuously updated.