Anthropic's Project Pilot Puts Claude Behind the Controls of a Surveillance Drone
Anthropic has extended its physical-world capability research into the air. Project Pilot, published July 24, is the third experiment in a series that started with an AI-operated vending machine (Project Vend) and continued with a robot fetching physical objects (Project Fetch). The new work tests whether Claude models can autonomously pilot a consumer drone through a locate-and-follow surveillance task — and introduces a dedicated benchmark, Drone-Bench, to track that capability over time.
The experiments were built with Andon Labs. In Drone-Bench, a model must command a drone using only its camera feed and sensor data to identify a moving target and maintain pursuit. It is the kind of task used in aerial surveillance, inspection, and tracking applications. Anthropic has not published specific pass rates, but the system card notes that Claude’s ability to use off-the-shelf robots is “on track to approach the ease with which coding agents use software tools” — the implication being that physical autonomy is catching up to the pace of digital agent capability.
Why Anthropic Is Publishing This
The framing throughout Project Pilot is explicitly safety-first. Anthropic’s Frontier Red Team ran these evaluations to establish situational awareness — understanding how close current models are to effective autonomous physical operation before that capability reaches general deployment.
That framing is deliberate. Drone technology has obvious dual-use implications: the same locate-and-follow behaviour that helps a photographer track a subject in a crowded stadium can be repurposed for targeted surveillance or threat engagement. By measuring where Claude sits on that capability curve internally and publishing the methodology, Anthropic is establishing a public record before the capability matures in ways that are harder to track.
Project Fetch Phase 2 findings, referenced in the paper, already showed improvement fast enough that Anthropic considers it “approaching” the ease of software tool use. Project Pilot extends that measurement to a more consequential physical domain.
What Drone-Bench Measures
Drone-Bench is structured around a single core task: locate a target and follow it. The evaluation runs in real hardware with an off-the-shelf drone, not a simulated environment. The model receives video and sensor input and must generate flight commands in a closed-loop system.
Locating and following is a relatively bounded task — the drone does not need to reason about broader mission objectives or coordinate with other systems. Anthropic chose this scope deliberately: it is complex enough to be a meaningful capability signal but narrow enough to evaluate cleanly. A model that cannot reliably execute locate-and-follow is not a credible threat for more complex autonomous physical missions.
The benchmark joins a growing set of physical-world capability evaluations that Anthropic has developed internally, alongside software-domain benchmarks like SWE-bench and Terminal-Bench. The pattern of publishing these before the capability is fully mature, rather than after, is consistent with Anthropic’s stated model of responsible capability disclosure.
What Comes Next
Project Fetch Phase 2 noted that models are approaching the point where robotic tool use becomes as reliable as digital tool use. If that observation holds, Drone-Bench scores will move quickly. The gap between a research evaluation and a deployable autonomous system in this domain is narrow in ways it is not for, say, scientific discovery tasks. That proximity is precisely why Anthropic is tracking it now.