GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
← Back to feed

China's Z.AI Releases GLM-5.1, Tops SWE-Bench Pro Ahead of GPT-5.4 and Claude Opus 4.6

A Chinese model is now best in the world at the toughest software engineering benchmark.

Z.AI, the Beijing-based lab formerly known as Zhipu AI, released GLM-5.1 this week. Its headline number: 58.4 on SWE-Bench Pro, clearing GPT-5.4 (57.7), Claude Opus 4.6 (57.3), and Gemini 3.1 Pro (54.2). It is the first Chinese model to top the SWE-Bench Pro leaderboard, and it does so while running on zero Nvidia hardware.

Key Numbers

  • SWE-Bench Pro: 58.4 — #1 overall, 0.7 pts ahead of GPT-5.4
  • NL2Repo: 42.7 — top score, tests ability to generate entire repository structures
  • Architecture: 744B-parameter Mixture-of-Experts, 40B active per token, 200K context window
  • Hardware: Huawei Ascend chips — no Nvidia dependency

What It Is

GLM-5.1 is a post-training upgrade to GLM-5 — same architecture, same context window, but with a reinforcement learning pipeline retargeted specifically at coding distributions. The base GLM-5 was already the first open model to score 50+ on the Artificial Analysis Intelligence Index. GLM-5.1 pushes further.

Z.AI claims the model can run autonomously for up to eight hours, refining strategies across thousands of iterations — what the company calls “long-horizon agentic engineering.”

Why It Matters

In a field where frontier models are separated by fractions of a point, a full point lead on SWE-Bench Pro is not noise. It signals that Chinese labs are no longer catching up to US frontier models on coding — they’re ahead on at least one benchmark that matters.

The Nvidia-free architecture is the other story. Built on Huawei Ascend chips, GLM-5.1 demonstrates that frontier-level coding capability is achievable entirely outside the US export control perimeter. That has implications beyond benchmark tables.