Anthropic, Amazon, Microsoft, and Google Propose a CVSS for AI Jailbreaks
Anthropic published a draft AI jailbreak severity framework on July 2, timed to coincide with Fable 5’s global return. The framework is the first attempt by a frontier lab to create shared terminology for describing how dangerous a given jailbreak actually is — and it is backed by Glasswing partners including Amazon, Microsoft, and Google.
The problem it addresses is real: jailbreaks vary enormously in what they unlock. Some bypass minor behavioural guardrails. Others allow a model to generate working malware or provide step-by-step cyberattack instructions. Right now, when labs or governments talk about a jailbreak, there is no agreed-upon severity scale. Every disclosure uses its own language.
The Four-Tier Framework
Anthropic’s proposed system classifies cybersecurity-related AI behaviors into four categories:
| Category | Intended Classifier Behavior |
|---|---|
| Prohibited use | Block — ransomware, C2 infrastructure, malware dev, cyber-physical sabotage, defense evasion |
| High-risk dual use | Block — widely used by attackers, limited defensive utility |
| Low-risk dual use | Monitor, sometimes block — mostly defensive, some attacker value |
| Benign use | Allow with monitoring |
The prohibited tier is specific. It includes ransomware, wipers, malware delivery, command-and-control servers, defense evasion (AV/EDR bypass, log tampering), cyber-physical sabotage of power and water infrastructure, and internet backbone attacks including BGP hijacking and DNS root manipulation.
The low-risk dual-use category is where the framework gets useful. This tier overlaps with what Anthropic calls the “safety margin” — benign requests that are blocked out of caution because they are indistinguishable from higher-risk requests at classifier time. Fable 5 runs with a larger safety margin than previous models, meaning more false positives in exchange for greater confidence in catching harmful behaviors.
Why Now
Fable 5’s government-ordered suspension in mid-June was triggered by a jailbreak. That event exposed the absence of a shared language: Anthropic, regulators, and industry partners had no common framework for describing what the jailbreak actually unlocked or how severe the resulting capability was. The framework is Anthropic’s attempt to build that vocabulary before the next incident.
The classification system is written as a starting point, not a final standard. Anthropic is soliciting feedback from academia, civil society, and government via cyber-safeguards@anthropic.com, with the explicit goal of turning the draft into an agreed-upon industry standard over time.
HackerOne Program
Alongside the framework, Anthropic launched a HackerOne program for security researchers to submit potential cyber jailbreaks in Fable 5. This is a meaningful operational step — it creates a structured intake path for the kind of vulnerability that triggered the export control order, and routes findings to Anthropic’s safety team rather than into ad-hoc disclosure or public exploit threads.
The framework document is at anthropic.com/news/fable-safeguards-jailbreak-framework. The HackerOne program is live at hackerone.com/anthropic-cyber-jailbreak.