GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 —
GROK-46H 865 —
GEM-37FH 865 —
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 —
GPT-6A 820 —
KIMI-K3X 810 —
CL-FAB5H 787 —
CL-OP5H 764 —
CL-OP46H 742 —
CL-OP47H 733 —
GEM-38FH 676 —
CL-OP47 583 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

Claude Fable 5.1 System Card: Chartography Hits 86.2% With Tools, Autonomy Risk Holds at Low

Anthropic published the Claude Fable 5.1 and Claude Mythos 5.1 system card on September 1, 2026, adding a new professional data analysis benchmark to its standard evaluation suite and affirming that neither model crosses any safety threshold from its August 2026 risk report.

Chartography benchmark

The headline new evaluation is Chartography — a benchmark for professional data analysis tasks developed with Surge AI. Scores with tools enabled:

ModelWithout toolsWith tools
Claude Fable 5.142.6%86.2%
Claude Fable 536.6%84.2%
Claude Opus 529.6%83.0%

Fable 5.1 gains 5.6 percentage points over Fable 5 on the no-tools baseline, a larger relative improvement than the 2-point gap in the tool-enabled condition. That spread suggests the 5.1 update strengthened base analytical reasoning rather than just tool-call efficiency.

The without-tools gap between Opus 5 (29.6%) and Fable 5.1 (42.6%) is 13 points — notable given that Opus 5 leads Fable 5.1 on SWE-bench Verified (97.0% vs ~95.0%). Chartography and SWE-bench measure different skills: Chartography is grounded in structured data analysis with domain-expert-verified answer ranges, while SWE-bench tests code-level debugging. The two rankings don’t directly contradict each other.

Autonomy and safety assessment

Anthropic’s August 2026 risk report assessed two autonomy threat models for Fable 5.1:

  • Threat model 1 (autonomous capability misuse): classified as “low” risk, consistent with the prior report. Anthropic cites specific mitigations detailed in the August report.
  • Threat model 2 (AI R&D acceleration): assessed as not applicable to Mythos 5.1. The system card states Mythos 5.1 has AI R&D capabilities comparable to the frontier set by Mythos 5, and does not cross the risk threshold under the same two-factor reasoning applied to Mythos 5.

The structure of the assessment is notable: rather than releasing a per-model safety report, Anthropic now publishes one comprehensive risk report covering all models at a point in time (the August report covers models and actions as of July 15, 2026), then references that report in each subsequent system card. The system card itself is model-specific capability documentation, not a new risk analysis.

What changed in 5.1

The system card positions Fable 5.1 as a targeted update rather than a full retrain. The Mythos 5.1 comparison with Mythos 5 on AI R&D capabilities — described as “comparable” — implies the safety-relevant capability boundaries didn’t move significantly. The Chartography improvement and the 25% price reduction published in prior reports (Terminal-Bench-Science: 52.6%, per Stack Futures’ earlier coverage) represent the primary differentiation over Fable 5.

For operators evaluating whether to migrate from Fable 5 to 5.1, the system card provides the clearest official signal that the risk profile hasn’t changed while the capability floor on professional analytical tasks has risen.