GPT-5.6 Sol xHigh Codex-Harness Enters Code Arena at #2 — 13 Points Behind Fable 5
OpenAI’s GPT-5.6 Sol xHigh, evaluated using its internal Codex harness, entered Code Arena on July 10 at 1636 ELO — second only to Claude Fable 5’s 1649. It is the first non-Anthropic model to clear 1620 on the leaderboard, and the first time any OpenAI model has finished within 15 points of the Code Arena summit.
The Leaderboard
| Rank | Model | ELO | Votes |
|---|---|---|---|
| 1 | Claude Fable 5 | 1649 | 2,161 |
| 2 | GPT-5.6 Sol xHigh (codex-harness) | 1636 | 624 |
| 3 | GLM-5.2 (max) | 1580 | 4,290 |
| 4 | Grok 4.5 | 1566 | 1,598 |
| 5 | Claude Opus 4.8 Thinking | 1560 | 6,710 |
| 6 | Claude Opus 4.7 Thinking | 1557 | 9,932 |
| 18 | GPT-5.5 xHigh (codex-harness) | 1502 | 8,601 |
The jump from GPT-5.5 codex-harness (1502, rank 18) to GPT-5.6 Sol codex-harness (1636, rank 2) is 134 ELO points. That is the largest single-generation gain recorded on Code Arena — and it moved OpenAI’s harness submission from outside the top 15 to second place.
The Gap at the Top
Fable 5 leads by 13 points. The Sol codex-harness submission has only 624 votes against Fable 5’s 2,161 and GLM-5.2’s 4,290, so the confidence interval on that gap is still wide. With 8,601 votes, the GPT-5.5 codex-harness submission is the most-voted OpenAI model on the board; Sol is still accumulating.
What the data already shows is structural. Below the top two, the next model — GLM-5.2 at 1580 — sits 56 ELO below Sol and 69 below Fable 5. The frontier has compressed to a two-model cluster, with Anthropic and OpenAI holding both positions. Third place is not competitive with either.
What the Codex Harness Means
OpenAI’s Codex harness is the same agent scaffold used in Codex CLI and the backend of ChatGPT Work’s coding mode. It handles tool routing, context management, and multi-turn repair loops. The harness submission to Code Arena is not a raw model evaluation — it is OpenAI’s production coding agent, benchmarked against Anthropic’s production coding model.
The comparison is direct. Fable 5 holds Code Arena #1 with the same architecture that powers Claude Code and Claude Managed Agents. Sol xHigh with Codex is OpenAI’s equivalent stack. The 13-point gap is the competitive distance between the two labs’ full coding pipelines as of July 10.
Previous Generation Baseline
The GPT-5.5 codex-harness submission, with 8,601 votes, sits at 1502 — rank 18 on the same leaderboard. The 134-point generation gain means GPT-5.6 Sol did not just improve on the underlying model; the combination of model quality and harness tuning together closed a gap that had persisted across multiple prior update cycles.
For context, the ELO distance between GPT-5.5 codex-harness (1502) and the current #1 (1649) would have made it appear structurally uncompetitive. The upgrade erased that distance entirely except for the final 13 points.
The vote count for the Sol submission will grow over the coming weeks. Whether the 13-point gap to Fable 5 holds, narrows, or reverses is the open question.