GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Mythos Field Report: US Banks Rush to Patch, cURL Developer Says It Found One Bug

Two data points on Claude Mythos, published within 48 hours of each other, tell a contradictory story about what the model actually does in the wild.

The Bank Case

Reuters reported on May 12 that US banks — including major institutions — are conducting urgent IT remediation programs based on vulnerabilities Mythos surfaced. The model, which Anthropic has restricted to vetted security teams, was used to scan internal systems and produced findings that prompted software upgrades and emergency patching cycles.

Anthropic’s commercial framing for Mythos positions it as a force multiplier for defensive security: it can chain multiple steps of analysis — code review, exploit feasibility assessment, patch drafting — that would normally require a senior security engineer weeks to complete. Reuters’ reporting suggests at least some financial institutions are treating it that way, with urgency.

The cURL Case

Daniel Stenberg, creator of cURL — one of the most audited open-source projects in existence — was given access to Mythos and pointed it at his project. Result: one vulnerability.

Stenberg published his assessment: the hype around Mythos is “primarily marketing.” His argument is straightforward. cURL has been continuously reviewed by human security researchers for decades. A model finding a single issue in that environment says less about the model’s capability than it does about the project’s existing security posture. The test was a ceiling test, not a floor test.

Why Both Are Probably Right

The discrepancy is not a contradiction — it is a signal about deployment context. Financial institutions running large legacy codebases, internal services with limited prior security review, and proprietary systems that have never seen a formal red team are exactly the environment where an AI model with Mythos-class capabilities would find genuine issues at scale. A well-maintained open-source project with continuous community review is not.

Anthropic’s positioning of Mythos as a security tool has always been predicated on the claim that most enterprise infrastructure has never been systematically reviewed. The bank results are consistent with that claim. Stenberg’s result is consistent with an exception to it.

The Hype Problem

The NYT framing — “Is Anthropic’s New A.I. Really That Scary? It Depends Whom You Ask” — captures the marketing tension Anthropic is navigating. Mythos launched with deliberate danger-framing: too capable to release publicly. That framing serves commercial purposes (scarcity, prestige, regulatory negotiating leverage) but creates an expectation mismatch when developers with well-audited projects run their own tests.

The model is real. The findings at financial institutions are real. The danger framing is also doing work that isn’t purely about danger.