GPT-56T 861 —
MUSE-SPK 837 +0.2%
GPT-56SC 790 -4.6%
GLM-5 781 -0.4%
CL-OP55X 780 -5.1%
GROK-46H 780 -5.1%
QWEN-38X 748 -9.2%
GPT-6A 743 -9.4%
KIMI-K3X 742 —
CL-FAB5H 698 -6.1%
CL-OP5H 675 -6.2%
GEM-38FH 672 -0.7%
CL-OP5X 670 -5.5%
CL-OP55H 668 —
CL-OP46H 657 -5.9%
CL-OP47H 648 -6.1%
GPT-56S 618 -0.6%
GEM-37FH 610 -7.2%
GEM-36FH 593 —
CL-OP48H 588 —
CL-OP47 581 -0.2%
GEM-35FH 580 —
GPT-55H 541 -7%
INKL 531 —
GEM-31P 512 -0.2%
CL-OP46 498 +0.4%
GEM-3P 498 -0.2%
CL-OP48 492 +0.4%
GPT-52 464 —
GPT-55 423 —
GPT-56T 861 —
MUSE-SPK 837 +0.2%
GPT-56SC 790 -4.6%
GLM-5 781 -0.4%
CL-OP55X 780 -5.1%
GROK-46H 780 -5.1%
QWEN-38X 748 -9.2%
GPT-6A 743 -9.4%
KIMI-K3X 742 —
CL-FAB5H 698 -6.1%
CL-OP5H 675 -6.2%
GEM-38FH 672 -0.7%
CL-OP5X 670 -5.5%
CL-OP55H 668 —
CL-OP46H 657 -5.9%
CL-OP47H 648 -6.1%
GPT-56S 618 -0.6%
GEM-37FH 610 -7.2%
GEM-36FH 593 —
CL-OP48H 588 —
CL-OP47 581 -0.2%
GEM-35FH 580 —
GPT-55H 541 -7%
INKL 531 —
GEM-31P 512 -0.2%
CL-OP46 498 +0.4%
GEM-3P 498 -0.2%
CL-OP48 492 +0.4%
GPT-52 464 —
GPT-55 423 —
← Back to feed

OpenAI Codex + GPT-5.5 Takes Terminal-Bench 2.0 at 82.0% — First Time OpenAI's Own Agent Leads

GPT-5.5 entered the Terminal-Bench 2.0 leaderboard on April 23 and immediately claimed the top position. Paired with OpenAI’s Codex agent, it scores 82.0% (±2.2), displacing ForgeCode’s 81.8% run with GPT-5.4 that had held the lead since mid-March.

The Current Top 5

RankAgentModelDateScore
1CodexGPT-5.5Apr 23, 202682.0% ±2.2
2ForgeCodeGPT-5.4Mar 12, 202681.8% ±2.0
3TongAgentsGemini 3.1 ProMar 13, 202680.2% ±2.6
4ForgeCodeClaude Opus 4.6Mar 12, 202679.8% ±1.6
5SageAgentGPT-5.3-CodexMar 13, 202678.4% ±2.2

Two Stories in One Data Point

The raw number is 0.2 points better. But the leaderboard change carries a structural argument: OpenAI’s in-house Codex agent, previously limited to OpenAI-flavored tooling, now outscores ForgeCode — the third-party scaffold that previously extracted the best results from GPT-5.4.

That’s a reversal from the March pattern, where ForgeCode outperformed Simple Codex (OpenAI’s baseline) by 6.7 points on the same model. OpenAI appears to have closed the internal scaffold gap, at least for GPT-5.5.

Context on GPT-5.5

GPT-5.5 launched with 82.7% Terminal-Bench self-reported and 58.6% SWE-Bench Pro. The April 23 independent submission on Terminal-Bench 2.0 comes in at 82.0% — a modest 0.7-point haircut from the launch figure, which is within the confidence interval (±2.2). That alignment between claimed and independently reproduced scores is better than average for frontier model launches.

Where the Ceiling Is

The 82.0% Terminal-Bench score is now 1.8 points ahead of TongAgents+Gemini 3.1 Pro, and 2.2 points ahead of ForgeCode+Claude Opus 4.6. Given confidence intervals, GPT-5.5+Codex and ForgeCode+GPT-5.4 are statistically overlapping — but every other model is now clearly behind.

The open question: whether any lab submits a newer model or new scaffold combination before the next generation of releases. Claude Opus 4.7 (launched April 16) has no Terminal-Bench submission yet. That entry, when it comes, will determine whether 82% is a ceiling or a waypoint.