GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
GLM-52 897 —
GPT-56SC 873 —
CL-OP5X 865 -0.9%
GROK-46H 865 -0.9%
GEM-37FH 865 -0.9%
GPT-56T 861 —
GLM-5 856 —
MUSE-SPK 841 —
QWEN-38X 824 -2.3%
GPT-6A 820 —
KIMI-K3X 810 -1%
CL-FAB5H 787 -0.9%
CL-OP5H 764 -0.9%
CL-OP46H 742 -0.9%
CL-OP47H 733 -1.1%
GEM-38FH 676 -1%
CL-OP47 585 -0.7%
INKL 531 —
CL-OP46 496 -0.2%
CL-OP48 490 -0.2%
← Back to feed

GPT-5.6 Sol xHigh Codex-Harness Enters Code Arena at #2 — 13 Points Behind Fable 5

OpenAI’s GPT-5.6 Sol xHigh, evaluated using its internal Codex harness, entered Code Arena on July 10 at 1636 ELO — second only to Claude Fable 5’s 1649. It is the first non-Anthropic model to clear 1620 on the leaderboard, and the first time any OpenAI model has finished within 15 points of the Code Arena summit.

The Leaderboard

RankModelELOVotes
1Claude Fable 516492,161
2GPT-5.6 Sol xHigh (codex-harness)1636624
3GLM-5.2 (max)15804,290
4Grok 4.515661,598
5Claude Opus 4.8 Thinking15606,710
6Claude Opus 4.7 Thinking15579,932
18GPT-5.5 xHigh (codex-harness)15028,601

The jump from GPT-5.5 codex-harness (1502, rank 18) to GPT-5.6 Sol codex-harness (1636, rank 2) is 134 ELO points. That is the largest single-generation gain recorded on Code Arena — and it moved OpenAI’s harness submission from outside the top 15 to second place.

The Gap at the Top

Fable 5 leads by 13 points. The Sol codex-harness submission has only 624 votes against Fable 5’s 2,161 and GLM-5.2’s 4,290, so the confidence interval on that gap is still wide. With 8,601 votes, the GPT-5.5 codex-harness submission is the most-voted OpenAI model on the board; Sol is still accumulating.

What the data already shows is structural. Below the top two, the next model — GLM-5.2 at 1580 — sits 56 ELO below Sol and 69 below Fable 5. The frontier has compressed to a two-model cluster, with Anthropic and OpenAI holding both positions. Third place is not competitive with either.

What the Codex Harness Means

OpenAI’s Codex harness is the same agent scaffold used in Codex CLI and the backend of ChatGPT Work’s coding mode. It handles tool routing, context management, and multi-turn repair loops. The harness submission to Code Arena is not a raw model evaluation — it is OpenAI’s production coding agent, benchmarked against Anthropic’s production coding model.

The comparison is direct. Fable 5 holds Code Arena #1 with the same architecture that powers Claude Code and Claude Managed Agents. Sol xHigh with Codex is OpenAI’s equivalent stack. The 13-point gap is the competitive distance between the two labs’ full coding pipelines as of July 10.

Previous Generation Baseline

The GPT-5.5 codex-harness submission, with 8,601 votes, sits at 1502 — rank 18 on the same leaderboard. The 134-point generation gain means GPT-5.6 Sol did not just improve on the underlying model; the combination of model quality and harness tuning together closed a gap that had persisted across multiple prior update cycles.

For context, the ELO distance between GPT-5.5 codex-harness (1502) and the current #1 (1649) would have made it appear structurally uncompetitive. The upgrade erased that distance entirely except for the final 13 points.

The vote count for the Sol submission will grow over the coming weeks. Whether the 13-point gap to Fable 5 holds, narrows, or reverses is the open question.