GLM 5.2 Debuts at #10 on Agent Arena — Z.ai's Open-Weight Model Cracks the Proprietary Tier
Z.ai’s GLM 5.2 (Max) was added to the Agent Arena leaderboard on June 16 and debuted at rank 10, placing it inside a tier previously occupied entirely by closed proprietary models from Anthropic, OpenAI, and Google.
Its composite score of 4.37% (±2.48%) puts it above Claude Opus 4.8 standard (rank 11, 3.60%) and Claude Sonnet 4.6 (rank 12, 3.22%). The comparison is imperfect — GLM 5.2 has only 9,625 votes, roughly one-third the vote count of the established models, which explains the wider uncertainty margin. Position will stabilise as more battles accumulate.
The context: GLM 5.1 sits at rank 13 with 2.66%. GLM 5.2’s debut at 10 represents a meaningful generational jump for the Z.ai model family.
Category Breakdown
The composite conceals some category-level variance worth noting:
| Category | GLM 5.2 (Max) | Claude Opus 4.8 Thinking |
|---|---|---|
| Coding | 9.43% | 10.75% |
| Web | 14.88% | 14.36% |
| Computer Use | 6.00% | 10.48% |
| Research | 1.69% | 9.08% |
| Agent Safety | 1.86% | 0.56% |
Web tasks are GLM 5.2’s strongest category (14.88%), where it edges Claude Opus 4.8 Thinking (14.36%). Research tasks (1.69%) represent its weakest showing — well below the Anthropic and OpenAI frontier.
MiniMax M3 Also Enters
MiniMax M3 was added to the Agent Arena on the same day and debuted at rank 19 with a 2.79% composite (±1.70%), 10,258 votes. It places below GLM 5.2 and above Qwen 3.6 Plus (rank 20, 4.24% with 32,121 votes). Given MiniMax M3’s higher vote count, Qwen 3.6 Plus’s lower position is likely more reliable at this sample size.
Open-Weight Position
With GLM 5.2 at rank 10, Z.ai holds two positions in the top 13 (GLM 5.1 at rank 13). Both models carry MIT licenses and route through SiliconFlow. The Agent Arena top 10 is now:
- Claude Fable 5 (High) — 14.17%
- Claude Opus 4.8 (Thinking) — 9.04%
- GPT-5.5 (xHigh) — 8.27% 4–5. Claude Opus 4.7 / 4.7 Thinking — ~8.1%
- GPT-5.5 (High) — 7.78%
- GPT-5.5 — 6.73%
- Claude Opus 4.6 — 6.73%
- GPT-5.4 (High) — 6.54%
- GLM 5.2 (Max) — 4.37% (MIT, open-weight)
A 2.17 point gap separates GPT-5.4 High (#9) from GLM 5.2 (#10). That gap will narrow or widen as vote counts equalise. The ranking is live and continuously updated.