vix Scaffold Puts Claude Opus 4.7 at 90.2% on Terminal-Bench 2.0 — 5.5 Points Clear of GPT-5.5
The Terminal-Bench 2.0 leaderboard reset today. The vix scaffold — built by china-qijizhifeng and running on Claude Opus 4.7 — posted 90.2% accuracy (±2.1), the first entry to cross 90% on the benchmark and 5.5 points ahead of the next GPT-5.5-powered entry.
Current Top-10
| Rank | Agent | Model | Accuracy | Date |
|---|---|---|---|---|
| 1 | vix | Claude Opus 4.7 | 90.2% ±2.1 | 2026-05-15 |
| 2 | NexAU-AHE | GPT-5.5 | 84.7% ±2.1 | 2026-05-14 |
| 3 | LemonHarness | Multiple | 84.5% ±2.6 | 2026-05-14 |
| 4 | Capy | GPT-5.5 | 83.1% ±2.1 | 2026-05-14 |
| 5 | Polaris | Multiple | 82.2% ±2.8 | 2026-05-14 |
| 6 | Codex CLI | GPT-5.5 | 82.0% ±2.2 | 2026-04-23 |
| 7 | ForgeCode | GPT-5.4 | 81.8% ±2.0 | 2026-03-12 |
| 8 | WOZCODE | Claude Opus 4.7 | 80.2% ±2.1 | 2026-05-14 |
What Changed
OpenAI’s Codex CLI held the #1 spot at 82.0% since late April — the first time OpenAI’s first-party scaffold had topped the leaderboard. That lead lasted three weeks. vix erased it by 8.2 points in a single submission.
Three GPT-5.5-powered entries (NexAU-AHE, Capy, Codex CLI) cluster between 82.0% and 84.7%. The performance ceiling for GPT-5.5 on Terminal-Bench 2.0 appears to be in that range regardless of scaffolding. Opus 4.7 has now broken past it with a 5.5-point margin.
Anthropic’s own WOZCODE scaffold running on Opus 4.7 sits at rank 8 with 80.2%. The gap between WOZCODE and vix — 10 points, same model — underscores how much scaffolding strategy influences results at the frontier. Benchmark performance is increasingly a systems engineering problem, not just a base model problem.
Scoring Context
Terminal-Bench 2.0 tests agents against real terminal tasks: file manipulation, system administration, shell scripting, and multi-step environment navigation. The benchmark runs agents in a minimal bash environment with no special tooling — the same mini-SWE-agent-style loop used for the SWE-bench “Bash Only” track. Results reflect base model capability plus scaffold efficiency under constrained tooling.
At 90.2%, Opus 4.7 via vix now leads every major agentic coding benchmark where Anthropic models have submitted: SWE-bench Verified (87.6%, per Anthropic’s launch figures), Terminal-Bench 2.0 (90.2%), and SWE-bench Pro (where Opus 4.7 held a 5.7-point lead over GPT-5.5 as of April).