GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

NVIDIA Nemotron 3 Ultra Enters Agent Arena — US Open-Weights Gets Its First Real-World Agent Test

NVIDIA Nemotron 3 Ultra was added to Arena’s Agent Arena leaderboard on June 12, one day after GPT-5.5 (xHigh) claimed second place in the same ranking. The entry places the US’s top open-weights model into its first real-world agentic evaluation — a benchmark measured not by static coding tasks but by live user sessions testing task success, error recovery, steerability, tool fidelity, and satisfaction across multi-step agent interactions.

What the Model Brings In

Nemotron 3 Ultra scores 48 on the Artificial Analysis Intelligence Index, which makes it the strongest US-origin open-weights model available today and the closest American competitor to the Chinese open-weights cluster sitting at 52-54. It operates at 550 billion total parameters with 90% sparsity, activating roughly 55 billion per forward pass. On pre-release DeepInfra infrastructure, it served over 300 tokens per second — throughput that its Chinese-origin competitors at similar intelligence levels rarely match.

What it has never been tested on is agent performance under real conditions: multi-turn tasks with tool calls, mid-session failures, and unsatisfied users who abandon partial completions. That is what Agent Arena provides.

The Field It Enters

Agent Arena currently fields 11 models with statistical confidence. Claude Fable 5 (High) leads at 12.94% composite, supported by 14,177 votes and a ±2.10% confidence band that separates it cleanly from second place. The next four positions — GPT-5.5 (xHigh) at 10.56%, Claude Opus 4.8 (Thinking) at 9.26%, Claude Opus 4.7 (Thinking) at 8.64%, and GPT-5.5 (High) at 8.16% — are all proprietary frontier models from Anthropic or OpenAI. Anthropic holds seven of the top eleven spots.

No open-weights model has made the top-11 table yet. Nemotron 3 Ultra is the first US-origin open model to accumulate votes in the live leaderboard. It will take several thousand sessions before its composite score stabilizes into a number worth quoting.

The Test That Matters

The AA Intelligence Index 48 score measures reasoning, coding, and instruction-following on benchmarks. It does not measure what happens when a model drops a tool call, receives corrective input from a user, and needs to recover without halting the task. That gap — between benchmark performance and agent completion rates — is what Agent Arena’s architecture is designed to surface.

Fable 5’s widest margin over the field on Agent Arena is not in task success (17.55% vs Opus 4.8’s 9.88%) but in error recovery (28.86% vs GPT-5.5 High’s 11.07%). The divide between benchmark and agent scores has appeared before: Opus 4.7’s SWE-bench numbers did not translate proportionally to an Agent Arena lead over GPT-5.4. Whether a 48-point intelligence-index model performs closer to its rating or underperforms it on agents is not predictable from intelligence benchmarks alone.

Open-Weights in the Agent Era

The Agent Arena result for Nemotron 3 Ultra will be one of the first data points on whether open-weights models can approach proprietary model agent performance. The current benchmark story for open-weights is mixed: on SWE-bench Verified, Kimi K2.6 (80.2%) and DeepSeek V4 Pro (80.6%) sit close to proprietary competitors. But SWE-bench measures single-session code repair; Agent Arena’s composite spans sessions with unsatisfied users and interrupted workflows.

NVIDIA’s Nemotron 3 Ultra enters the ranking as a speed-efficient open model in a field still dominated by closed-access systems. Its position on the table — when enough votes accumulate to generate a stable estimate — will be watched by enterprise teams building on open-weights infrastructure who need an agent capability number, not just a reasoning index score.