GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Booz Allen: Claude Mythos Scored 80 on Cyber Weapon Index, Only Model to Complete Full Kill Chain — But a Rank-15 Model Matched It With a Harness

Booz Allen Hamilton released its inaugural Cyber Weapon Index (CWI), the most comprehensive head-to-head evaluation of AI offensive cyber capability published by a major defence contractor. Eighteen frontier models — nine American, nine Chinese — were run autonomously against a production-grade enterprise network with no curated tool menus or additional scaffolding. Every action was validated through network telemetry, host logs, domain controller data, and intrusion-detection sensors.

Claude Mythos, Anthropic’s most capable model, scored 80 on the CWI and is the only system to have completed the full cyber kill chain autonomously. No other model came close on the raw index.

Full scores

ModelCWI Score
Claude Mythos80
Grok-4.549
GPT-5.6 Sol46
Muse Spark 1.138
Kimi K338
GLM-5.237
Claude Opus 4.836
GPT-5.5-Cyber34
Nemotron-Ultra33

Mythos achieved the full kill chain — reconnaissance, initial access, foothold, credential access, privilege escalation, lateral movement, and domain compromise — both with and without stolen employee credentials. Without credentials it still ultimately compromised the domain, adapting its strategy dynamically rather than following a scripted attack path.

Four models reached full domain access and control. Another four achieved lateral movement. Two progressed through credential access. Only one model failed to penetrate the network at all.

The finding that matters most

The headline number is not the most significant part of the report. Booz Allen separately tested what happens when a lower-ranked model is given a commodity attack harness — scaffolding software that wraps the model with a structured tool loop and memory. A model ranked 15th of 18 matched Mythos on domain access.

The implication: evaluating models in isolation materially understates real-world offensive cyber risk. The unit of danger is not the model. It is the model plus harness plus tools plus autonomy.

This is the same finding that Terminal-Bench’s multi-harness tier has been surfacing for months — the right scaffold can close a 30-point benchmark gap. In cyber contexts that gap maps to whether a threat actor gets into a network.

No US-China capability gap

Booz Allen found no substantial differentiation between US and Chinese models. Both camps produced models capable of lateral movement and domain access. The report treats this as a more alarming baseline than the Mythos-leads-by-31-points framing suggests: near-frontier offensive AI is already globally distributed.

The firm assessed that most of the 18 models will reach full kill chain capability within six months, placing the window for precautionary infrastructure hardening in early 2027 at the latest.

What Booz Allen is recommending

Three prescriptions in the report: machine-speed defences (AI defending at AI attack speed), binding readiness standards for critical infrastructure, and continuous measurement of global AI capability. The last point is an implicit critique of point-in-time evaluations that miss capability proliferation between annual assessments.

The report was published August 17, 2026. It is available as a public PDF at boozallen.com.