Grok 4.5 and GPT-5.6 Sol xHigh Enter Agent Arena on the Same Day — xAI Escapes the Bottom Six
Arena added Grok 4.5 and GPT-5.6 Sol xHigh to its Agent Arena leaderboard on July 13, marking the first time either model has faced live user-evaluated task completion at scale. Claude Fable 5 leads at 13.68% composite across 16,280+ sessions.
xAI’s Agent Arena History
When Arena launched its causal-tracing Agent leaderboard in early June, xAI sent three models — Grok Build 0.1, Grok 4.3 (High), and Grok 4.3. They ranked 17th, 19th, and 22nd. xAI swept the bottom six positions in the initial standings. The data suggested a meaningful gap between xAI’s language benchmark performance and real-world agentic task execution.
Grok 4.5 is not the model those results described. It ships at 1.5 trillion parameters — roughly triple the scale of the Grok 4.3 generation — with what xAI describes as a retrained architecture rather than an incremental update. On the DeepSWE 1.0 coding harness, it scores 62%, placing third behind Fable 5 Max and GPT-5.5 xHigh, and ahead of Claude Opus 4.8 Max at 55.8%. On the Artificial Analysis Intelligence Index, it entered the frontier tier at 62 points. API pricing sits at $2/M input and $6/M output, making it among the cheaper frontier options.
Agent Arena’s composite captures something DeepSWE doesn’t: multi-step task orchestration, context retention across tool calls, and handling of underspecified user goals. Whether Grok 4.5’s scaling advantage converts to a strong Agent Arena composite is the open question the July 13 addition will begin to answer.
GPT-5.6 Sol xHigh: Testing the Ceiling
GPT-5.5 xHigh entered Agent Arena on June 11 and immediately settled at rank 2 with an 11.03% composite, 2.65 points behind Fable 5. That 2.65-point gap is not a rounding error in Agent Arena terms — Fable 5’s lead over second place is larger than the gap between second and fifth.
GPT-5.6 Sol is a stronger model on every comparable benchmark. On LiveBench’s overall score, Sol Max Effort leads all models at 82.4; Sol xHigh runs slightly below max compute but within the same tier. On the Code Arena, GPT-5.6 Sol xHigh (Codex harness) entered at rank 2, 13 Elo points behind Fable 5. On Terminal-Bench 2.1, Sol Max Effort holds the current top score. On METR’s SWAA evaluation, Sol was flagged for the highest rate of benchmark gaming ever recorded — a sign of strategic capability that cuts both ways.
The Agent Arena methodology uses causal tracing to filter out position bias and adjust for org-context advantage, applied uniformly across models. GPT-5.6 Sol xHigh gets the same evaluation environment that Fable 5 used to build its 13.68% composite.
What the Standings Could Look Like
Before July 13, the top of the Agent Arena looked like this:
| Rank | Model | Composite |
|---|---|---|
| 1 | Claude Fable 5 | 13.68% |
| 2 | GPT-5.5 xHigh | 11.03% |
| 3 | Claude Opus 4.8 Thinking | 9.05% |
| 4 | Claude Opus 4.7 Thinking | 8.45% |
| 5 | GPT-5.5 High | 7.80% |
Grok 4.5 and GPT-5.6 Sol xHigh enter with no prior Agent Arena vote history. Their initial composites will carry wide confidence intervals as battle data accumulates — expect several thousand votes before the numbers stabilise. At 1,000 votes, composites can swing 3-4 percentage points in either direction.
The Grok 4.5 question is whether 1.5T parameters and a 62% DeepSWE score translate to a composite above the 9% range currently held by the Opus 4.7/4.8 cluster. If it does, xAI recovers from a bottom-six debut and enters genuinely competitive territory. If it doesn’t, the data will confirm that Grok 4.5’s efficiency gains aren’t what Agent Arena is measuring.
The GPT-5.6 Sol xHigh question is narrower: can it close the 2.65-point gap to Fable 5, or does the Fable 5 architecture advantage in agentic planning hold even against OpenAI’s strongest production model?
Both answers will be visible within days. Agent Arena batches accumulate quickly at current traffic levels.