GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Claude Fable 5.1 Sets New Benchmark Ceiling: 52.6% Terminal-Bench-Science, 25% Cheaper Than Fable 5

Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 today. The two models share the same weights. The difference is safeguards: Fable 5.1 is generally available via API, AWS, Google Cloud, and Azure under model ID claude-fable-5-1. Mythos 5.1 is restricted to vetted organizations through the Cyber Verification Program and a new biology access program developed in partnership with the US government.

Benchmark Results

Fable 5.1 tops Anthropic’s published comparison across six benchmark suites:

BenchmarkFable 5.1Fable 5Opus 5GPT-5.6 Sol
Terminal-Bench-Science 0.152.6%24.7%29.0%22.4%
Terminal-Bench 4.055.8%42.0%—37.3%
Terminal-Bench 4.0 (Mythos 5.1)60.9%———
GDPval-AA v21853182317111711
OSWorld 2.0 (partial)77.9%75.4%72.9%—
AutomationBench31.4%17.1%26.9%19.6%
CursorBench 3.2.073.4%70.5%70.0%67.2%
HLE (with tools)65.0%—63.6%—

The Terminal-Bench-Science 0.1 score is the standout: Fable 5.1 at 52.6% nearly doubles Opus 5’s 29.0%, on a benchmark specifically designed to measure autonomous scientific research capability.

Anthropic notes evaluation caveats: where production safeguards intervened, Fable 5.1 and Fable 5 scored zero on OSWorld 2.0 tasks, and biology tasks were completed by Opus 5 as fallback. The scores reflect production-condition behavior, not an uncapped model.

Pricing and Access Changes

Fable 5.1 costs approximately 25% less than Fable 5 for token-billed workloads. The reduction comes entirely from cache read pricing. For highly agentic pipelines that make heavy use of cached context, Anthropic estimates savings up to 45%.

Two other changes address enterprise friction:

Enterprise Frontier Safeguards (EFS): A new architecture that provides zero-data-retention privacy by storing data in customer-controlled cloud infrastructure rather than Anthropic’s. EFS rolls out to enterprise customers in phases starting later this fall. Until then, eligible customers can use Fable 5.1 with zero data retention under existing terms.

Reduced false positives in cybersecurity: Fable 5.1 can now assist with discovering software vulnerabilities (but not developing exploits). The updated safeguards flag 60% fewer false positives than Fable 5 on cybersecurity tasks.

Scientific Research Demonstrations

Anthropic ran two verified scientific demonstrations with Mythos 5.1 and Fable 5.1:

Protein binders: Mythos 5.1 designed protein binders, sent to two external organizations for experimental validation. On three targets, the designs outperformed the best submissions to Adaptyv Bio’s protein design competitions. The hit rate — the fraction of designs that were viable binders — reached approximately 50% across 12 targets. Industry baseline is 10-15%.

Venus elevation mapping: Fable 5.1 trained a neural network on 30-year-old NASA Magellan radar imagery to produce a new elevation map covering one-third of Venus. The map resolves surface details at 2-3 km versus the prior 10-20 km, and improves height accuracy by up to 25%. Anthropic released the map under Creative Commons ahead of the NASA VERITAS and ESA EnVision missions.

Both demonstrations used external tool access. Anthropic is positioning these as early evidence of AI’s role in scientific discovery, not proof of general autonomous research capability.

Deployment

Fable 5.1 defaults to High effort in Claude Code, Medium effort in Claude Cowork and Claude.ai. Anthropic notes that at Low or Medium effort, Fable 5.1 matches or exceeds Fable 5’s High-effort scores at lower cost — meaning the default settings at launch are not extracting full performance, by design.

Mythos 5.1’s biology access program will open enrollment for scientists soon, per Anthropic.