GPT-6 Astra Scores 99.9% on ARC-AGI-3, Leaving Every Competitor Below 31%
GPT-6 Astra scored 99.9% on ARC-AGI-3, according to results posted Thursday by the ARC Prize organization. Every other model that has run on the benchmark sits below 31%.
That gap is the headline, and it is larger than the model’s other credentials warrant alone. On ARC-AGI-1 and ARC-AGI-2, the frontier has converged: Astra, GPT-5.6 Sol, Claude Fable 5.1, Claude Fable 5, and Claude Opus 5 all cluster between 89% and 99% on ARC-AGI-2. No single model dominates. On ARC-AGI-3, Astra is alone.
What ARC-AGI-3 Tests
ARC Prize describes ARC-AGI-3 as measuring four components of agentic intelligence. Unlike the earlier iterations — which tested abstract visual pattern completion — ARC-AGI-3 places models in interactive environments that require tool use, planning across steps, and recovery from unexpected states. The task is not to recognize a pattern from examples but to operate effectively in an unfamiliar setting.
The benchmark also introduces an efficiency dimension absent from prior versions. Humans completing ARC-AGI-3 tasks take a certain number of actions per level; the system tracks whether AI models are faster or slower than the median tested human. Astra used fewer actions than the median human on 96% of levels — meaning it found more efficient solutions, not just correct ones.
ARC Prize notes one specific observed behavior: Astra built compact symbolic world models of unfamiliar environments. Rather than operating reactively, it encoded the rules of a novel setting into an internal representation, then reasoned from that model. This is the behavior alignment researchers have been most interested in detecting, and most worried about evaluating, in opaque architectures.
The Comparison Table
| Model | ARC-AGI-3 | ARC-AGI-2 | ARC-AGI-1 |
|---|---|---|---|
| GPT-6 Astra | 99.9% | 95.0% | 98.5% |
| Claude Opus 5 | 30.2% | 90.4% | 97.5% |
| GPT-5.6 Sol | 7.8% | 92.5% | 97.5% |
| Claude Fable 5.1 | — | 90.0% | 97.5% |
| Claude Fable 5 | — | 89.2% | 98.5% |
Claude Opus 5, the highest-scoring non-Astra model, is at 30.2% — a result that represents genuine agentic capability and is itself well above human baseline. The drop to GPT-5.6 Sol at 7.8% is steep. ARC-AGI-3 appears to be testing something the current generation of non-Astra models cannot reliably do.
OpenAI’s AGI Claim
OpenAI President Greg Brockman, at Thursday’s launch briefing, said Astra brings the frontier “fully into the AGI era.” The ARC-AGI-3 result is one pillar of that claim. The other is the cybersecurity Critical threshold designation — the first time OpenAI has classified a model as capable of finding and exploiting undiscovered vulnerabilities without step-by-step human guidance.
ARC Prize was founded specifically as an AGI evaluation, with the stated goal of measuring “the residual gap between current AI and AGI.” The definition is operational: AGI is a system that can acquire any skill a human can, as efficiently as a human can. Astra’s action efficiency result — more efficient than the median human on 96% of levels — is ARC Prize’s own formulation, not OpenAI’s marketing.
Whether that constitutes AGI depends on how much weight you put on the benchmark versus what is absent from it. ARC-AGI-3 tests novel environment navigation under controlled conditions. It does not test long-horizon planning in the real world, physical-world embodiment, sustained autonomous goal pursuit over days or weeks, or the ability to learn a new domain without any prior exposure to its conceptual vocabulary.
What the Gap Means for Rivals
The 99.9% versus 30.2% spread is a capability discontinuity. On reasoning, coding, and general knowledge benchmarks, the frontier has been tightly contested for the past year. Anthropic, Google, and xAI have all fielded models within a few percentage points of OpenAI on most evaluations.
ARC-AGI-3 is an exception. Either Astra’s recurrent depth architecture contributes something qualitatively different to novel environment reasoning, or the other labs have simply not run their top models at sufficient effort levels yet, or ARC-AGI-3 rewards something specific to Astra’s training. The ARC Prize leaderboard remains open — results from other models at other effort levels will either close the gap or confirm it.
Until that data arrives, Astra holds a benchmark position with no recent precedent.