GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Anthropic Built Opus 4.7 for Cyber Defense. It Scored Last.

Anthropic released Claude Opus 4.7 in April with explicit cybersecurity marketing: the Cyber Verification Program, automated detection of high-risk prompts, and product copy describing “cyber-grade safeguards.” It was the first Claude built for security work. On Simbian’s Cyber Defense Benchmark — 1,206 real attack-log investigations across 13 of 14 MITRE ATT&CK tactics and 105 chained kill-chain procedures — Opus 4.7 finished last.

The Leaderboard

RankModelCoverageMedian CostTactics Passing
1Claude Opus 4.644.5%$2.717 of 13
2Claude Sonnet 4.636.9%$2.212 of 13
3Claude Opus 4.832.7%$1.691 of 13
4Claude Opus 4.730.9%$1.650 of 13

GPT-5.5 scored 26.4%, GPT-5 scored 17.4%. Every Anthropic model on the board outranked every OpenAI and Google model. The best LLM for SOC work today is the one released five months before the cyber-specific one.

Why the Cyber Model Performs Worst

Simbian’s explanation is structural. Opus 4.7 optimizes for speed and cost: its median investigation runs $1.65 at 243 seconds — the fastest and cheapest Claude tested. For cybersecurity defense work, that is a liability. SOC reasoning requires hypothesis generation under heavy noise, MITRE-coverage planning, and iterative queries across 100K+ events per environment. Investigation depth tracks investigation time. The model designed to move fast through security tasks moves too fast to do security work well.

The Cyber Verification Program safeguards block offensive prompts as designed. They do not add defensive reasoning skill. The benchmark measures defense, not safety posture, and those are not the same thing.

Opus 4.8 does not recover the gap. At 32.7% coverage, it passes 1 of 13 MITRE tactics and trails Opus 4.6 by 12 points. Newer models optimize for general benchmarks. The defensive reasoning problem that matters for SOC work — noise, hypothesis stacking, long sessions — runs orthogonal to those benchmarks.

The Harder Finding: No Model Crosses 50%

Across 14 frontier models from Anthropic, OpenAI, Google, and open-weight providers, none crossed 50% coverage. Frontier LLMs routinely score above 80% on offensive security benchmarks. Defense is the structurally harder problem, and the gap between the two is growing wider as labs train for offense.

The practical upshot: raw LLMs should not be placed directly in front of a SOC queue. The benchmark tested the same scenarios with Simbian’s own harness wrapped around the best bare LLM — coverage went from 46% to 95%, independently verified by a global MSSP in April 2026. Forty-nine points came from context, skills, and agent loop design. One point came from the model choice.

What This Means for Anthropic’s Cyber Push

Anthropic has built real infrastructure around cybersecurity: Glasswing (200+ organizational partners, 10,000+ vulnerabilities patched), Claude Mythos (NSA access, classified deployments), and the Opus 4.7 Cyber Verification Program. The lab’s positioning in security is genuine. But the Simbian numbers surface a gap between offensive safeguard quality and defensive capability quality that marketing has not bridged.

The recommendation from the benchmark data is to run Opus 4.6 for SOC investigation work and wrap it in an agentic harness. Sonnet 4.6 is the price-performance pick at $2.21 per investigation. Skip Opus 4.7 for this specific workload.