China's Z.AI Releases GLM-5.1, Tops SWE-Bench Pro Ahead of GPT-5.4 and Claude Opus 4.6
A Chinese model is now best in the world at the toughest software engineering benchmark.
Z.AI, the Beijing-based lab formerly known as Zhipu AI, released GLM-5.1 this week. Its headline number: 58.4 on SWE-Bench Pro, clearing GPT-5.4 (57.7), Claude Opus 4.6 (57.3), and Gemini 3.1 Pro (54.2). It is the first Chinese model to top the SWE-Bench Pro leaderboard, and it does so while running on zero Nvidia hardware.
Key Numbers
- SWE-Bench Pro: 58.4 — #1 overall, 0.7 pts ahead of GPT-5.4
- NL2Repo: 42.7 — top score, tests ability to generate entire repository structures
- Architecture: 744B-parameter Mixture-of-Experts, 40B active per token, 200K context window
- Hardware: Huawei Ascend chips — no Nvidia dependency
What It Is
GLM-5.1 is a post-training upgrade to GLM-5 — same architecture, same context window, but with a reinforcement learning pipeline retargeted specifically at coding distributions. The base GLM-5 was already the first open model to score 50+ on the Artificial Analysis Intelligence Index. GLM-5.1 pushes further.
Z.AI claims the model can run autonomously for up to eight hours, refining strategies across thousands of iterations — what the company calls “long-horizon agentic engineering.”
Why It Matters
In a field where frontier models are separated by fractions of a point, a full point lead on SWE-Bench Pro is not noise. It signals that Chinese labs are no longer catching up to US frontier models on coding — they’re ahead on at least one benchmark that matters.
The Nvidia-free architecture is the other story. Built on Huawei Ascend chips, GLM-5.1 demonstrates that frontier-level coding capability is achievable entirely outside the US export control perimeter. That has implications beyond benchmark tables.