OpenAI Codex + GPT-5.5 Takes Terminal-Bench 2.0 at 82.0% — First Time OpenAI's Own Agent Leads
GPT-5.5 entered the Terminal-Bench 2.0 leaderboard on April 23 and immediately claimed the top position. Paired with OpenAI’s Codex agent, it scores 82.0% (±2.2), displacing ForgeCode’s 81.8% run with GPT-5.4 that had held the lead since mid-March.
The Current Top 5
| Rank | Agent | Model | Date | Score |
|---|---|---|---|---|
| 1 | Codex | GPT-5.5 | Apr 23, 2026 | 82.0% ±2.2 |
| 2 | ForgeCode | GPT-5.4 | Mar 12, 2026 | 81.8% ±2.0 |
| 3 | TongAgents | Gemini 3.1 Pro | Mar 13, 2026 | 80.2% ±2.6 |
| 4 | ForgeCode | Claude Opus 4.6 | Mar 12, 2026 | 79.8% ±1.6 |
| 5 | SageAgent | GPT-5.3-Codex | Mar 13, 2026 | 78.4% ±2.2 |
Two Stories in One Data Point
The raw number is 0.2 points better. But the leaderboard change carries a structural argument: OpenAI’s in-house Codex agent, previously limited to OpenAI-flavored tooling, now outscores ForgeCode — the third-party scaffold that previously extracted the best results from GPT-5.4.
That’s a reversal from the March pattern, where ForgeCode outperformed Simple Codex (OpenAI’s baseline) by 6.7 points on the same model. OpenAI appears to have closed the internal scaffold gap, at least for GPT-5.5.
Context on GPT-5.5
GPT-5.5 launched with 82.7% Terminal-Bench self-reported and 58.6% SWE-Bench Pro. The April 23 independent submission on Terminal-Bench 2.0 comes in at 82.0% — a modest 0.7-point haircut from the launch figure, which is within the confidence interval (±2.2). That alignment between claimed and independently reproduced scores is better than average for frontier model launches.
Where the Ceiling Is
The 82.0% Terminal-Bench score is now 1.8 points ahead of TongAgents+Gemini 3.1 Pro, and 2.2 points ahead of ForgeCode+Claude Opus 4.6. Given confidence intervals, GPT-5.5+Codex and ForgeCode+GPT-5.4 are statistically overlapping — but every other model is now clearly behind.
The open question: whether any lab submits a newer model or new scaffold combination before the next generation of releases. Claude Opus 4.7 (launched April 16) has no Terminal-Bench submission yet. That entry, when it comes, will determine whether 82% is a ceiling or a waypoint.