GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

vix Scaffold Puts Claude Opus 4.7 at 90.2% on Terminal-Bench 2.0 — 5.5 Points Clear of GPT-5.5

The Terminal-Bench 2.0 leaderboard reset today. The vix scaffold — built by china-qijizhifeng and running on Claude Opus 4.7 — posted 90.2% accuracy (±2.1), the first entry to cross 90% on the benchmark and 5.5 points ahead of the next GPT-5.5-powered entry.

Current Top-10

RankAgentModelAccuracyDate
1vixClaude Opus 4.790.2% ±2.12026-05-15
2NexAU-AHEGPT-5.584.7% ±2.12026-05-14
3LemonHarnessMultiple84.5% ±2.62026-05-14
4CapyGPT-5.583.1% ±2.12026-05-14
5PolarisMultiple82.2% ±2.82026-05-14
6Codex CLIGPT-5.582.0% ±2.22026-04-23
7ForgeCodeGPT-5.481.8% ±2.02026-03-12
8WOZCODEClaude Opus 4.780.2% ±2.12026-05-14

What Changed

OpenAI’s Codex CLI held the #1 spot at 82.0% since late April — the first time OpenAI’s first-party scaffold had topped the leaderboard. That lead lasted three weeks. vix erased it by 8.2 points in a single submission.

Three GPT-5.5-powered entries (NexAU-AHE, Capy, Codex CLI) cluster between 82.0% and 84.7%. The performance ceiling for GPT-5.5 on Terminal-Bench 2.0 appears to be in that range regardless of scaffolding. Opus 4.7 has now broken past it with a 5.5-point margin.

Anthropic’s own WOZCODE scaffold running on Opus 4.7 sits at rank 8 with 80.2%. The gap between WOZCODE and vix — 10 points, same model — underscores how much scaffolding strategy influences results at the frontier. Benchmark performance is increasingly a systems engineering problem, not just a base model problem.

Scoring Context

Terminal-Bench 2.0 tests agents against real terminal tasks: file manipulation, system administration, shell scripting, and multi-step environment navigation. The benchmark runs agents in a minimal bash environment with no special tooling — the same mini-SWE-agent-style loop used for the SWE-bench “Bash Only” track. Results reflect base model capability plus scaffold efficiency under constrained tooling.

At 90.2%, Opus 4.7 via vix now leads every major agentic coding benchmark where Anthropic models have submitted: SWE-bench Verified (87.6%, per Anthropic’s launch figures), Terminal-Bench 2.0 (90.2%), and SWE-bench Pro (where Opus 4.7 held a 5.7-point lead over GPT-5.5 as of April).