GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897
GPT-56SC 873
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861
GLM-5 856
MUSE-SPK 841
QWEN-38X 824 -2.3%
GPT-6A 820
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

AISI: Open-Weight Models Are Now 4-7 Months Behind Closed Frontier on Cyber

The UK AI Security Institute published its first public comparative analysis of open and closed frontier models on cyber capability. The finding: GLM-5.2, the most cyber-capable open-weight model at time of testing, now trails the closed frontier by four to seven months. In early 2025, that gap was six to ten months.

AISI used two evaluation methods to triangulate the result.

Narrow Cyber Tasks

The Narrow Cyber Tasks benchmark runs 70 tasks across four difficulty levels, covering vulnerability research, reverse engineering, web exploitation, and cryptography. GLM-5.2 (released June 2026) matched the performance of Claude Opus 4.6 (February 2026) — approximately four months behind the closed frontier. DeepSeek V4-Pro scored at roughly the Opus 4.5 level (November 2025), placing it about five months behind.

Cyber Ranges

The Cyber Ranges benchmark tests longer-horizon, multi-step attack scenarios that require sustained planning across many turns. On these, GLM-5.2 matched Claude Opus 4.5 (November 2025), putting the trailing distance at approximately seven months. DeepSeek V4-Pro performed at a similar level.

Both methods converge on the same directional conclusion: the gap was 6-10 months for most of 2025. It is 4-7 months now.

Compute as Multiplier

One finding reframes how to think about risk trajectory. Long-horizon cyber capability scales with inference compute. GPT-5.6 Sol completed all 32 steps of AISI’s most demanding evaluation, “The Last Ones,” in 7 of 10 attempts when given a 100M-token budget per run. Performance continued improving as the token budget increased.

This matters for open-weight models specifically. A model released at a fixed capability level can be made substantially more capable by anyone willing to spend more on inference — no retraining required. The capability embedded in open weights is not static.

The Window for Defenders

AISI frames the narrowing gap as a shrinking preparation window. When the most capable open-weight model matches closed frontier performance from four months ago, defenders with access to closed systems had four months to harden against those capabilities before they became available without controls.

Once an open-weight model is released, the controls available to closed providers disappear permanently. Safety guardrails can be removed, weights can be redistributed, and the model can run on private infrastructure beyond any monitoring. AISI calls this “a persistent and irreversible risk of misuse.”

The report does not recommend against open-weight release in general, but frames each release decision as requiring capability-level assessment — not just a binary open-versus-closed judgment.

Key Numbers

ModelBenchmarkMatched Closed EquivalentGap
GLM-5.2Narrow Cyber TasksOpus 4.6 (Feb 2026)~4 months
GLM-5.2Cyber RangesOpus 4.5 (Nov 2025)~7 months
DeepSeek V4-ProNarrow Cyber TasksOpus 4.5 (Nov 2025)~5 months
GPT-5.6 Sol”The Last Ones” (32-step)7/10 runs at 100M tokens

Gap trend: 6-10 months (early 2025) to 4-7 months (July 2026).