GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Anthropic, Amazon, Microsoft, and Google Propose a CVSS for AI Jailbreaks

Anthropic published a draft AI jailbreak severity framework on July 2, timed to coincide with Fable 5’s global return. The framework is the first attempt by a frontier lab to create shared terminology for describing how dangerous a given jailbreak actually is — and it is backed by Glasswing partners including Amazon, Microsoft, and Google.

The problem it addresses is real: jailbreaks vary enormously in what they unlock. Some bypass minor behavioural guardrails. Others allow a model to generate working malware or provide step-by-step cyberattack instructions. Right now, when labs or governments talk about a jailbreak, there is no agreed-upon severity scale. Every disclosure uses its own language.

The Four-Tier Framework

Anthropic’s proposed system classifies cybersecurity-related AI behaviors into four categories:

CategoryIntended Classifier Behavior
Prohibited useBlock — ransomware, C2 infrastructure, malware dev, cyber-physical sabotage, defense evasion
High-risk dual useBlock — widely used by attackers, limited defensive utility
Low-risk dual useMonitor, sometimes block — mostly defensive, some attacker value
Benign useAllow with monitoring

The prohibited tier is specific. It includes ransomware, wipers, malware delivery, command-and-control servers, defense evasion (AV/EDR bypass, log tampering), cyber-physical sabotage of power and water infrastructure, and internet backbone attacks including BGP hijacking and DNS root manipulation.

The low-risk dual-use category is where the framework gets useful. This tier overlaps with what Anthropic calls the “safety margin” — benign requests that are blocked out of caution because they are indistinguishable from higher-risk requests at classifier time. Fable 5 runs with a larger safety margin than previous models, meaning more false positives in exchange for greater confidence in catching harmful behaviors.

Why Now

Fable 5’s government-ordered suspension in mid-June was triggered by a jailbreak. That event exposed the absence of a shared language: Anthropic, regulators, and industry partners had no common framework for describing what the jailbreak actually unlocked or how severe the resulting capability was. The framework is Anthropic’s attempt to build that vocabulary before the next incident.

The classification system is written as a starting point, not a final standard. Anthropic is soliciting feedback from academia, civil society, and government via cyber-safeguards@anthropic.com, with the explicit goal of turning the draft into an agreed-upon industry standard over time.

HackerOne Program

Alongside the framework, Anthropic launched a HackerOne program for security researchers to submit potential cyber jailbreaks in Fable 5. This is a meaningful operational step — it creates a structured intake path for the kind of vulnerability that triggered the export control order, and routes findings to Anthropic’s safety team rather than into ad-hoc disclosure or public exploit threads.

The framework document is at anthropic.com/news/fable-safeguards-jailbreak-framework. The HackerOne program is live at hackerone.com/anthropic-cyber-jailbreak.