Anthropic's Red Team: Mythos Preview Geolocates Outdoor Photos to 37 km, Writes Working Drone Software
Anthropic’s Frontier Red Team published a research paper Wednesday measuring what current frontier and open-weight models can do on two tasks with direct military applications: locating people from outdoor photographs, and writing software to guide drone weapons through complex environments.
The evaluations are intended to track AI progress on dual-use capabilities as a leading indicator of risk — similar to how the lab tracks biological capability uplift. Both tasks showed consistent gains across model generations, and the researchers conclude that the patterns are not specific to Claude: they are a challenge for every lab and for policymakers.
Geolocation: 6,000 Outdoor Photos, Ranked Against Human Divisions
The intelligence targeting eval asked models to identify where 6,000 outdoor photos were taken. The benchmark used synthetic data with a tunable difficulty level — raising operational security of the subjects makes it harder for any system to correlate signals.
Results by model, sorted by median distance error:
| Model | Median error | Within 1 km |
|---|---|---|
| Mythos Preview | 37.0 km | 23.7% |
| Mythos 5 | 47.2 km | 23.1% |
| Opus 5 | 181 km | 18.0% |
| Sonnet 5 | 384 km | 9.9% |
| Kimi K3 (open-weight) | ~Sonnet-level median | 16.7% |
Mythos Preview and Mythos 5 score above Anthropic’s internal Master Division human baseline on both metrics. Opus 5 is roughly level with Master Division players. Sonnet 5 and Kimi K3 fall between expert and casual human tiers.
The Kimi K3 result is notable: its median distance error is at Sonnet’s level, but its within-1-km rate is approximately 7 times higher. The model places fewer photos at extreme distances than Sonnet even when its median is similar, suggesting a different error distribution rather than uniformly better performance.
Open-weights models generally score between Sonnet and the Mythos family. Perfect geolocation of outdoor photos is specifically described as “approaching superhuman” at the frontier.
Drone Guidance: Three Tasks, Full Coverage at the Frontier
The weapons development eval tested three software engineering tasks related to drone weapons: payload delivery guidance, terminal guidance (homing in on a moving target), and navigation through GPS interference. Each was run at easy, medium, and hard settings.
Opus 5, Mythos 5, and Mythos Preview wrote working guidance, navigation, and control software for every simulated task at all three difficulty levels. All three could iterate code into reliable performance on easier and medium settings.
Sonnet 5 completed the simplest version of each task and no more. Kimi K3 scored above Sonnet on payload delivery but fell to Sonnet’s level on terminal guidance and GPS interference — the two tasks requiring more precise real-time control logic.
What Anthropic Did in Response
The paper prompted Anthropic’s Safeguards team to implement new classifiers specifically targeting requests related to weapons development. The company acknowledges these classifiers will be imperfect because the underlying engineering capability is broadly useful — code that guides a drone also resembles code for autonomous vehicles and industrial robotics. The decision to deploy anyway reflects a preference for partial mitigation over none.
Anthropic explicitly says the results are not a Claude-specific problem. Open-weight models — including Kimi K3, which is publicly available — already show concerning capability levels on both tasks. Safety measures applied only to proprietary APIs do not address the full risk surface.