GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Opus 4.8 Thinking Enters Agent Arena at Rank 2 — Grok Build Lands at 17, xAI Sweeps the Bottom Six

Arena added four models to its Agent Arena leaderboard between June 8 and June 9 — Claude Opus 4.8 (Thinking), Claude Opus 4.8 (standard), Grok Build 0.1, and Grok 4.3 (High) — and the results are not what xAI’s positioning suggests.

The Updated Standings

RankModelComposite ScoreVotes
1GPT-5.5 (High)9.11%29,039
2Claude Opus 4.8 (Thinking) ↑ NEW9.06%19,691
3Claude Opus 4.7 (Thinking)8.42%27,154
4Claude Opus 4.68.15%27,299
5GPT-5.4 (High)7.97%28,950
6GPT-5.57.94%29,303
7Claude Opus 4.77.25%27,334
8Claude Opus 4.8 ↑ NEW4.35%17,328
…………
17Grok Build 0.1 ↑ NEW5.15%19,046
19Grok 4.3 (High) ↑ NEW9.23%7,168
22Grok 4.322.66%28,307

Agent Arena measures five real-world signals across hundreds of thousands of live user sessions in Arena’s Agent Mode: task success, steerability, error recovery, user praise versus complaints, and tool hallucination rate. The composite score uses a causal inference methodology that isolates each model’s contribution across these dimensions. A higher composite score is better.

The Thinking Mode Divide Is Wide

The same underlying model, Opus 4.8, produces dramatically different results depending on whether thinking is enabled. With thinking, it enters at rank 2 — 0.05 percentage points behind the leader. Without thinking, it drops to rank 8 at 4.35%. That is a 4.7-point gap between the same model in two operating modes.

The pattern is consistent across the Anthropic lineup: Claude Opus 4.7 (Thinking) ranks third at 8.42%, while Claude Opus 4.7 without thinking ranks seventh at 7.25%. Extended reasoning is not just a benchmark optimisation — it changes how an agent handles errors, corrections, and multi-step decisions under real task pressure.

Grok Build’s Reality Check

xAI launched Grok Build in May as a Claude Code competitor, citing a vendor-reported 70.8% on SWE-bench Verified and pricing $1 per million input tokens versus Claude Code’s higher rates. On the controlled benchmark, those are plausible credentials.

On Agent Arena, Grok Build 0.1 enters at rank 17 out of 22 models. The five-signal composite puts it at 5.15% — well below the top-tier cluster and below several models that launched months earlier. The gap between its SWE-bench number and its Agent Arena position reflects the difference between solving isolated GitHub issues and handling the full loop of multi-step tasks, user corrections, and error recovery under real conditions.

Grok 4.3 (High) entered at rank 19 with only 7,168 votes, so its 9.23% composite carries a wide confidence interval (±1.95%) and will firm up as more sessions accumulate. The base Grok 4.3 remains at rank 22 — last — with 28,307 votes and an 80.61% error recovery failure rate in the relevant sub-dimension, consistent with what Arena’s own analysis described as a terminal loop failure mode.

The Structural Finding

Six of the top seven positions on Agent Arena are held by thinking-enabled models or the standard GPT-5.5 High. The bottom six include three xAI entries. Grok Build was positioned as the cheaper alternative to Claude Code; the Agent Arena data suggests the capability gap it needs to close is not primarily in code generation, where SWE-bench measures it, but in the orchestration and recovery layer that defines whether an agent actually finishes jobs.

The leaderboard will continue updating as vote counts grow. Grok 4.3 (High)‘s rank is still forming. But the first data point for Grok Build’s real-world agentic performance is rank 17.