GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 —
GPT-56SC 789 -0.1%
GLM-5 781 —
CL-OP55X 779 -0.1%
GROK-46H 779 -0.1%
QWEN-38X 748 —
GPT-6A 743 —
KIMI-K3X 742 —
CL-FAB5H 697 -0.1%
CL-OP5H 674 -0.1%
GEM-38FH 672 —
CL-OP5X 669 -0.1%
CL-OP55H 667 -0.1%
CL-OP46H 656 -0.2%
CL-OP47H 647 -0.2%
GPT-56S 617 -0.2%
GEM-37FH 609 -0.2%
GEM-36FH 592 -0.2%
CL-OP48H 587 -0.2%
CL-OP47 580 -0.2%
GEM-35FH 579 -0.2%
GPT-55H 540 -0.2%
INKL 531 —
GEM-31P 511 -0.2%
CL-OP46 498 —
GEM-3P 498 —
CL-OP48 492 —
GPT-52 464 —
GPT-55 423 —
← Back to feed

Anthropic's Claude Mythos Preview Posts 93.9% on SWE-Bench Verified — a Generation Ahead of Opus 4.6

Anthropic has shared benchmark results for Claude Mythos Preview — an unreleased frontier model at the center of Project Glasswing, the company’s cybersecurity initiative. The numbers suggest a meaningful capability step over the current public generation.

Coding: The Numbers

BenchmarkMythos PreviewOpus 4.6Gap
SWE-bench Pro77.8%53.4%+24.4 pts
SWE-bench Verified93.9%80.8%+13.1 pts
SWE-bench Multilingual87.3%77.8%+9.5 pts
Terminal-Bench 2.082.0%65.4%+16.6 pts
SWE-bench Multimodal59.0%27.1%+31.9 pts

The SWE-bench Multimodal result is the most striking — Mythos scores more than double Opus 4.6 on a benchmark that tests AI ability to understand visual context in code (UI screenshots, diagrams, visual debugging).

When Gemini 3.1 Pro launched, GPT-5.3-Codex led SWE-bench Pro at 56.8%. Mythos Preview exceeds that by more than 21 points.

Reasoning

On GPQA Diamond — graduate-level science questions in physics, chemistry, and biology — Mythos Preview scores 94.6% versus 91.3% for Opus 4.6. The delta is smaller than on coding, but GPQA Diamond has been one of the harder benchmarks to move significantly.

Context

Mythos Preview is tied to Project Glasswing, Anthropic’s cybersecurity-focused initiative. The model is not publicly available. Anthropic has not announced pricing or a release timeline. These results are internal benchmarks shared by Anthropic — independent verification has not yet been published.