NVIDIA AVO Scores 100% on ARC-AGI-3 as the Harness Emerges as the Real AI Moat
ARC-AGI-3 was built to resist saturation. NVIDIA’s AVO architecture just scored 100% on it.
The result, published on NVIDIA’s technical blog this week, is significant not because of the number but because of what produced it. AVO is not a new model. It is a general-purpose agent architecture built around persistent state, tool use, grounded feedback, recovery mechanisms, and long-horizon context management. The model powering it is a component of the system, not the system itself.
Adel El Hallack, vice president of product in NVIDIA’s AI unit, made the distinction precise in comments to TechCrunch: “Generally speaking the world interprets an agent almost as an API of the model. But an agent is actually more than that. It is the model. It is the scaffolding around the model, which we call the harness.”
The argument that scaffolding beats raw model capability is not new. What is new is a leaderboard result at this level to back it.
What ARC-AGI-3 Measures
ARC-AGI-3 is a generalisation benchmark designed to test novel pattern recognition rather than memorisation of training examples. Achieving 100% on it with a harness-first approach signals that architectural design choices can now match or exceed gains from scaling model weights alone.
This matters for the competitive landscape in a specific way: model weights are increasingly commoditised and open-sourced. The harness is not. Oracle’s production database work illustrates the point independently. In their agent deployment, agents see typed inputs for fixed query templates and cannot write SQL at all. The constraint architecture, not the model capability, is what makes the system reliable enough for production.
The Moat Question
If the harness is the product, the companies building proprietary scaffolding are building the actual moat. NVIDIA’s AVO architecture demonstrates this at benchmark level. The companies that ship harness-plus-model as a unified system rather than an interchangeable API layer have a structural advantage that model-swapping does not eliminate.
The broader competitive implication is straightforward. Hyperscalers and lab incumbents are competing on model quality. The companies quietly winning on harness quality are accumulating an advantage that neither a weight update nor an API price cut will erase.
ARC-AGI-3 at 100% is one data point. But it is the kind of data point that tends to clarify a thesis that practitioners have been operating on for the last 18 months.