GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Terminal-Bench 3.0 Begins Development as TB2.0 Tops Saturate — and a Science Track Is Being Built Alongside

Terminal-Bench 3.0 is in active development. The benchmark team opened a contribution call in May targeting a specific design brief: harder tasks, longer horizons, richer terminal environments, and specialized domain expertise that TB2.0 cannot adequately test. The merge window closed at the end of May. Submitted tasks are now in build.

Alongside TB3.0, a separate track — terminal-bench-science — is also in development. The split design separates general terminal-environment agents from agents working in scientific computing contexts: code execution in scientific Python, numerical workflows, domain-specialized toolchains.

Why Now

The timing tracks directly with what the TB2.0 leaderboard is showing. The top five submissions from May 14 — NexAU-AHE at 84.7%, LemonHarness at 84.5%, Capy at 83.1%, Codex CLI at 82.2%, Polaris at 82.2% — sit within 2.5 percentage points. All confidence intervals overlap. At this compression, TB2.0 can no longer distinguish between leading-edge agents.

That is the same pattern playing out on SWE-bench Verified: Claude Fable 5 at 95.0%, Claude Mythos at 95.5%, Claude Opus 4.8 and GPT-5.5 both at 88.6-88.7%. The normalization ceiling for SWE-bench was set at 81% — a threshold that is now obsolete for every frontier model. The benchmark is measuring the gap between second tier and frontier, not between frontier models themselves.

What TB3.0 Is Targeting

The contribution rubric framed three difficulty axes:

  • Task horizon: multi-step sequences requiring maintained state across longer trajectories
  • Environment richness: runtime environments with more complexity than the bash shell at TB2.0’s core
  • Domain expertise: tasks requiring specialized knowledge that a general-purpose model cannot solve by pattern-matching on training data

The third axis is where terminal-bench-science diverges. Scientific computing tasks — running numerical simulations, debugging scientific Python pipelines, validating experiment outputs — constitute a distinct capability tier from system administration, networking, or code engineering tasks. Treating them as a separate benchmark rather than folding them into TB3.0 reflects what the benchmark community has learned from building TB2.0: domain mixing flattens the signal.

The Benchmark Shelf Life Problem

The benchmark development cycle now lags model capability by roughly 12-18 months. TB2.0 was published as frontier models were in the 55-70% range; by the time NexAU-AHE and LemonHarness reached 84-85%, the benchmark was already losing discrimination power at the top tier. TB3.0’s difficulty design is being calibrated against models that don’t yet exist — which means it needs to anticipate what GPT-5.5-successor and Claude Opus 4.9-class models will be capable of, not just what current models can’t solve.

LiveBench addresses this differently — releasing new questions continuously to prevent contamination. TB3.0’s approach is structural difficulty rather than freshness: tasks designed to be intrinsically harder to saturate, not just harder to memorize.

The result, if TB3.0 launches on schedule, is a benchmark ecosystem where the terminal environment tier has a meaningful top at some point past 2027 rather than the cluster compression now visible in TB2.0.