Mythos Field Report: US Banks Rush to Patch, cURL Developer Says It Found One Bug
Two data points on Claude Mythos, published within 48 hours of each other, tell a contradictory story about what the model actually does in the wild.
The Bank Case
Reuters reported on May 12 that US banks — including major institutions — are conducting urgent IT remediation programs based on vulnerabilities Mythos surfaced. The model, which Anthropic has restricted to vetted security teams, was used to scan internal systems and produced findings that prompted software upgrades and emergency patching cycles.
Anthropic’s commercial framing for Mythos positions it as a force multiplier for defensive security: it can chain multiple steps of analysis — code review, exploit feasibility assessment, patch drafting — that would normally require a senior security engineer weeks to complete. Reuters’ reporting suggests at least some financial institutions are treating it that way, with urgency.
The cURL Case
Daniel Stenberg, creator of cURL — one of the most audited open-source projects in existence — was given access to Mythos and pointed it at his project. Result: one vulnerability.
Stenberg published his assessment: the hype around Mythos is “primarily marketing.” His argument is straightforward. cURL has been continuously reviewed by human security researchers for decades. A model finding a single issue in that environment says less about the model’s capability than it does about the project’s existing security posture. The test was a ceiling test, not a floor test.
Why Both Are Probably Right
The discrepancy is not a contradiction — it is a signal about deployment context. Financial institutions running large legacy codebases, internal services with limited prior security review, and proprietary systems that have never seen a formal red team are exactly the environment where an AI model with Mythos-class capabilities would find genuine issues at scale. A well-maintained open-source project with continuous community review is not.
Anthropic’s positioning of Mythos as a security tool has always been predicated on the claim that most enterprise infrastructure has never been systematically reviewed. The bank results are consistent with that claim. Stenberg’s result is consistent with an exception to it.
The Hype Problem
The NYT framing — “Is Anthropic’s New A.I. Really That Scary? It Depends Whom You Ask” — captures the marketing tension Anthropic is navigating. Mythos launched with deliberate danger-framing: too capable to release publicly. That framing serves commercial purposes (scarcity, prestige, regulatory negotiating leverage) but creates an expectation mismatch when developers with well-audited projects run their own tests.
The model is real. The findings at financial institutions are real. The danger framing is also doing work that isn’t purely about danger.