GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Ox Alpha Posts 80% on DeepSWE: 15 Points Ahead of Fable 5, Zhipu AI Suspected

Developer Ben Davis ran a 10-task DeepSWE evaluation across several frontier models on the same harness and published the results on August 22. Ox Alpha, the anonymous model appearing on OpenRouter as stealth/ox-alpha, resolved 8 of 10 tasks for an 80% pass rate. The next closest was Claude Fable 5 Max at 65%.

The Numbers

ModelDeepSWE Pass Rate
Ox Alpha80% (8/10)
Claude Fable 5 Max65%
GLM-5.3 Max62%
Grok 4.6 xhigh62%
GPT-5.6 Sol Max52%

A second independent run on a separate DeepSWE subset returned 63% for Ox Alpha. That is lower, but still ahead of most top-tier models in that set. The gap between runs reflects different task selection and configuration rather than instability.

DeepSWE tasks require reading full code repositories, locating bugs, applying patches, and passing automated test suites: close to real software engineering work, not contrived puzzles.

Identity

The August 21 debut of Ox Alpha placed it in the free tier on both OpenRouter (stealth/ox-alpha) and OpenCode (x-preview-f-free), with a 1,048,576-token context window and tool-calling support. Developers noticed the model processes video tokens identically to Zhipu AI’s GLM-5V-Turbo and shows a consistent 75-millisecond time-to-first-token pattern that matches Zhipu’s infrastructure fingerprint. No lab has confirmed authorship.

The GLM-5.3 Max comparison score of 62% is notable: if Ox Alpha is built on GLM architecture, it would represent a significant internal advancement over the publicly released GLM-5.3 variant, not a marginal improvement.

Context

The 80% figure comes from a 10-task sample, which carries substantial variance. An 8/10 score could be 9/10 or 7/10 on a different draw. DeepSWE harness conditions, including repository access, timeout, and tool scaffolding, are standardized across this test but differ from SWE-bench Verified conditions. Direct cross-benchmark comparisons should be treated as directional.

What the numbers do establish: Ox Alpha can sustain multi-step agentic coding work long enough to close real software engineering tasks at a rate that challenges models from OpenAI, Anthropic, and xAI. If Zhipu AI is the source, the free-tier pricing makes it the highest-capability coding model available at zero marginal cost, by a significant margin.