Anthropic's Claude Mythos Preview Posts 93.9% on SWE-Bench Verified — a Generation Ahead of Opus 4.6
Anthropic has shared benchmark results for Claude Mythos Preview — an unreleased frontier model at the center of Project Glasswing, the company’s cybersecurity initiative. The numbers suggest a meaningful capability step over the current public generation.
Coding: The Numbers
| Benchmark | Mythos Preview | Opus 4.6 | Gap |
|---|---|---|---|
| SWE-bench Pro | 77.8% | 53.4% | +24.4 pts |
| SWE-bench Verified | 93.9% | 80.8% | +13.1 pts |
| SWE-bench Multilingual | 87.3% | 77.8% | +9.5 pts |
| Terminal-Bench 2.0 | 82.0% | 65.4% | +16.6 pts |
| SWE-bench Multimodal | 59.0% | 27.1% | +31.9 pts |
The SWE-bench Multimodal result is the most striking — Mythos scores more than double Opus 4.6 on a benchmark that tests AI ability to understand visual context in code (UI screenshots, diagrams, visual debugging).
When Gemini 3.1 Pro launched, GPT-5.3-Codex led SWE-bench Pro at 56.8%. Mythos Preview exceeds that by more than 21 points.
Reasoning
On GPQA Diamond — graduate-level science questions in physics, chemistry, and biology — Mythos Preview scores 94.6% versus 91.3% for Opus 4.6. The delta is smaller than on coding, but GPQA Diamond has been one of the harder benchmarks to move significantly.
Context
Mythos Preview is tied to Project Glasswing, Anthropic’s cybersecurity-focused initiative. The model is not publicly available. Anthropic has not announced pricing or a release timeline. These results are internal benchmarks shared by Anthropic — independent verification has not yet been published.