Claude Fable 5.1 Sets New Benchmark Ceiling: 52.6% Terminal-Bench-Science, 25% Cheaper Than Fable 5
Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 today. The two models share the same weights. The difference is safeguards: Fable 5.1 is generally available via API, AWS, Google Cloud, and Azure under model ID claude-fable-5-1. Mythos 5.1 is restricted to vetted organizations through the Cyber Verification Program and a new biology access program developed in partnership with the US government.
Benchmark Results
Fable 5.1 tops Anthropic’s published comparison across six benchmark suites:
| Benchmark | Fable 5.1 | Fable 5 | Opus 5 | GPT-5.6 Sol |
|---|---|---|---|---|
| Terminal-Bench-Science 0.1 | 52.6% | 24.7% | 29.0% | 22.4% |
| Terminal-Bench 4.0 | 55.8% | 42.0% | — | 37.3% |
| Terminal-Bench 4.0 (Mythos 5.1) | 60.9% | — | — | — |
| GDPval-AA v2 | 1853 | 1823 | 1711 | 1711 |
| OSWorld 2.0 (partial) | 77.9% | 75.4% | 72.9% | — |
| AutomationBench | 31.4% | 17.1% | 26.9% | 19.6% |
| CursorBench 3.2.0 | 73.4% | 70.5% | 70.0% | 67.2% |
| HLE (with tools) | 65.0% | — | 63.6% | — |
The Terminal-Bench-Science 0.1 score is the standout: Fable 5.1 at 52.6% nearly doubles Opus 5’s 29.0%, on a benchmark specifically designed to measure autonomous scientific research capability.
Anthropic notes evaluation caveats: where production safeguards intervened, Fable 5.1 and Fable 5 scored zero on OSWorld 2.0 tasks, and biology tasks were completed by Opus 5 as fallback. The scores reflect production-condition behavior, not an uncapped model.
Pricing and Access Changes
Fable 5.1 costs approximately 25% less than Fable 5 for token-billed workloads. The reduction comes entirely from cache read pricing. For highly agentic pipelines that make heavy use of cached context, Anthropic estimates savings up to 45%.
Two other changes address enterprise friction:
Enterprise Frontier Safeguards (EFS): A new architecture that provides zero-data-retention privacy by storing data in customer-controlled cloud infrastructure rather than Anthropic’s. EFS rolls out to enterprise customers in phases starting later this fall. Until then, eligible customers can use Fable 5.1 with zero data retention under existing terms.
Reduced false positives in cybersecurity: Fable 5.1 can now assist with discovering software vulnerabilities (but not developing exploits). The updated safeguards flag 60% fewer false positives than Fable 5 on cybersecurity tasks.
Scientific Research Demonstrations
Anthropic ran two verified scientific demonstrations with Mythos 5.1 and Fable 5.1:
Protein binders: Mythos 5.1 designed protein binders, sent to two external organizations for experimental validation. On three targets, the designs outperformed the best submissions to Adaptyv Bio’s protein design competitions. The hit rate — the fraction of designs that were viable binders — reached approximately 50% across 12 targets. Industry baseline is 10-15%.
Venus elevation mapping: Fable 5.1 trained a neural network on 30-year-old NASA Magellan radar imagery to produce a new elevation map covering one-third of Venus. The map resolves surface details at 2-3 km versus the prior 10-20 km, and improves height accuracy by up to 25%. Anthropic released the map under Creative Commons ahead of the NASA VERITAS and ESA EnVision missions.
Both demonstrations used external tool access. Anthropic is positioning these as early evidence of AI’s role in scientific discovery, not proof of general autonomous research capability.
Deployment
Fable 5.1 defaults to High effort in Claude Code, Medium effort in Claude Cowork and Claude.ai. Anthropic notes that at Low or Medium effort, Fable 5.1 matches or exceeds Fable 5’s High-effort scores at lower cost — meaning the default settings at launch are not extracting full performance, by design.
Mythos 5.1’s biology access program will open enrollment for scientists soon, per Anthropic.