GLM-5.3 Enters GDPval-AA at #3 with ELO 1769, Above Grok 4.6
Z AI’s GLM-5.3 (max) appeared on the GDPval-AA v2 leaderboard on August 21 at ELO 1769, placing third in the overall ranking. The full top four:
| Rank | Model | ELO |
|---|---|---|
| 1 | Claude Opus 5 (Max Effort) | 1845 |
| 2 | Claude Opus 5 (Xhigh Effort) | 1814 |
| 3 | GLM-5.3 (max) | 1769 |
| 4 | Grok 4.6 (high) | 1747 |
GLM-5.3 also landed on Arena.ai’s Text Arena and Code Arena leaderboards on August 19, two days before the GDPval-AA result.
What GLM-5.3’s Position Represents
GDPval-AA v2 is Artificial Analysis’s comprehensive agentic benchmark. The 22-point gap between GLM-5.3 and Grok 4.6 is significant at this tier — ELO differences at the top compress, and 22 points separates real performance levels.
GLM-5.2 reached the top of the open-weights Intelligence Index in June but remained a tier behind the proprietary leaders. GLM-5.3 crosses into a different bracket: it is now ahead of xAI’s Grok 4.6 and above all listed OpenAI and Google models on this leaderboard, which ranks by agentic task performance rather than a single-shot quality metric.
The distinction matters. GDPval-AA evaluates multi-turn agentic task completion, not exam-style benchmarks where frontier models are densely clustered. A third-place finish here is a stronger claim than a mid-table position on a general reasoning eval.
Z AI’s Trajectory
Z AI, the company behind the GLM series, released GLM-5.3 on August 19. Previous models in the GLM-5 family have consistently landed above their expected class given model size. GLM-5.2 matched GPT-5.5 on real-world agent tasks at a lower inference cost, according to Artificial Analysis data from June.
GLM-5.3’s GDPval-AA entry adds a third consecutive generation with competitive top-table positioning. The model is closed-weight and available via Z AI’s API.
DeepSeek-V4 Pro High, another Chinese-lab model, was also added to the Agent Arena and Text Arena leaderboards on August 19. Both additions land in the same week, continuing a pattern where Chinese labs clock multiple competitive model updates faster than Western labs can integrate benchmark results.