Anthropic Data: AI Task Horizons Doubling Every 4 Months, 8x Code Output — Global Pause Urged
Anthropic published a paper today called When AI Builds Itself that puts hard numbers on how fast AI is accelerating its own development. The data comes partly from internal Anthropic metrics that have not been reported before.
The headline figure: Anthropic engineers now ship 8x as much code per quarter as they did across the 2021–2025 period. The company attributes this entirely to AI coding agents taking over large portions of the software development cycle.
The Task Horizon Curve
The more consequential finding is the rate at which AI can autonomously complete long-duration tasks. The Anthropic Institute tracked this against a benchmark from METR — the organisation that measures how long AI can work independently on software tasks without human intervention.
The doubling rate has accelerated:
- Previous trend: task horizons doubling every 7 months
- Current trend: task horizons doubling every 4 months
The concrete checkpoints:
| Model | Date | Autonomous task duration |
|---|---|---|
| Claude Opus 3 | March 2024 | ~4 minutes |
| Claude Sonnet 3.7 | March 2025 | ~90 minutes |
| Claude Opus 4.6 | 2026 | ~12 hours |
METR separately found that Claude Mythos Preview could work for “at least 16 hours” — at the upper bound of what their measurement infrastructure can currently assess.
What the Projection Implies
If the 4-month doubling holds:
- Tasks taking a skilled human days could come into range this year
- Tasks taking a human weeks could be within reach by 2027
The Anthropic paper frames this explicitly: if AI systems can eventually build and train their own successors without human input, “recursive self-improvement” becomes operational. The paper states the company is not there yet and calls recursive self-improvement “not inevitable” — but flags it as potentially arriving before institutions are prepared.
The Policy Signal
The Wall Street Journal reports that Anthropic is using this paper to urge a global pause in AI development, or at minimum a coordinated international response before recursive capability becomes real.
This is a notable shift in posture. Anthropic has historically positioned itself as the safety-oriented lab that builds anyway to ensure the technology is developed responsibly. Publishing internal metrics that show compounding acceleration while simultaneously calling for restraint is a different kind of message.
The paper notes that the same capability trend that accelerates AI research also increases the stakes around containment, monitoring, and behavioural alignment — because if an AI can build its own successor, the usual human review steps compress or disappear.
What the Numbers Mean for the Ticker
The 8x code output figure applies across the entire Anthropic engineering organisation, not just to AI-specific work. If this pattern generalises across frontier labs — and there is no structural reason it would not — the acceleration in model capability has a built-in multiplier: each generation of models is developed faster than the last using the previous generation as tooling.
The implication for frontier model releases is more frequent major capability jumps with shorter gaps between them. The SFX-10 Index trajectory reflects this: the top-10 composite has moved significantly in the past six months and the rate of change is not slowing.