Terminal-Bench 3.0 Begins Development as TB2.0 Tops Saturate — and a Science Track Is Being Built Alongside
Terminal-Bench 3.0 is in active development. The benchmark team opened a contribution call in May targeting a specific design brief: harder tasks, longer horizons, richer terminal environments, and specialized domain expertise that TB2.0 cannot adequately test. The merge window closed at the end of May. Submitted tasks are now in build.
Alongside TB3.0, a separate track — terminal-bench-science — is also in development. The split design separates general terminal-environment agents from agents working in scientific computing contexts: code execution in scientific Python, numerical workflows, domain-specialized toolchains.
Why Now
The timing tracks directly with what the TB2.0 leaderboard is showing. The top five submissions from May 14 — NexAU-AHE at 84.7%, LemonHarness at 84.5%, Capy at 83.1%, Codex CLI at 82.2%, Polaris at 82.2% — sit within 2.5 percentage points. All confidence intervals overlap. At this compression, TB2.0 can no longer distinguish between leading-edge agents.
That is the same pattern playing out on SWE-bench Verified: Claude Fable 5 at 95.0%, Claude Mythos at 95.5%, Claude Opus 4.8 and GPT-5.5 both at 88.6-88.7%. The normalization ceiling for SWE-bench was set at 81% — a threshold that is now obsolete for every frontier model. The benchmark is measuring the gap between second tier and frontier, not between frontier models themselves.
What TB3.0 Is Targeting
The contribution rubric framed three difficulty axes:
- Task horizon: multi-step sequences requiring maintained state across longer trajectories
- Environment richness: runtime environments with more complexity than the bash shell at TB2.0’s core
- Domain expertise: tasks requiring specialized knowledge that a general-purpose model cannot solve by pattern-matching on training data
The third axis is where terminal-bench-science diverges. Scientific computing tasks — running numerical simulations, debugging scientific Python pipelines, validating experiment outputs — constitute a distinct capability tier from system administration, networking, or code engineering tasks. Treating them as a separate benchmark rather than folding them into TB3.0 reflects what the benchmark community has learned from building TB2.0: domain mixing flattens the signal.
The Benchmark Shelf Life Problem
The benchmark development cycle now lags model capability by roughly 12-18 months. TB2.0 was published as frontier models were in the 55-70% range; by the time NexAU-AHE and LemonHarness reached 84-85%, the benchmark was already losing discrimination power at the top tier. TB3.0’s difficulty design is being calibrated against models that don’t yet exist — which means it needs to anticipate what GPT-5.5-successor and Claude Opus 4.9-class models will be capable of, not just what current models can’t solve.
LiveBench addresses this differently — releasing new questions continuously to prevent contamination. TB3.0’s approach is structural difficulty rather than freshness: tasks designed to be intrinsically harder to saturate, not just harder to memorize.
The result, if TB3.0 launches on schedule, is a benchmark ecosystem where the terminal environment tier has a meaningful top at some point past 2027 rather than the cluster compression now visible in TB2.0.