GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

GPT-6 Astra Posts 57.9% on Terminal-Bench 4.0, 2.1 Points Ahead of Fable 5.1

OpenAI on September 10 published a work-oriented launch post for GPT-6 Astra that includes the first Terminal-Bench 4.0 comparison table with Fable 5.1 and GPT-5.6 Sol in the same row. GPT-6 Astra scores 57.9%, Claude Fable 5.1 scores 55.8%, and GPT-5.6 Sol scores 37.3%.

What Terminal-Bench 4.0 Measures

Terminal-Bench 4.0 is a harder successor to the 2.x series. Cognition’s SWE-2, which scored 92.8% on Terminal-Bench 2.1, scored 27.3% on Terminal-Bench 4.0 when independently tested — a 65-point gap that exposed TB2.x as undersaturated at the frontier. TB4.0’s harder task distribution means the current cluster of frontier models sits in the 27–58% range rather than the 80–92% range TB2.x produced.

The current TB4.0 snapshot across published scores:

ModelTB4.0 Score
GPT-6 Astra57.9%
Claude Fable 5.155.8%
GPT-5.6 Sol37.3%
Cognition SWE-227.3%

The 2.1-point gap between Astra and Fable 5.1 is within the margin of error for most agent benchmarks, which typically carry ±2% confidence intervals. OpenAI does not include error bars in its launch post.

Cost Framing

OpenAI’s comparison highlights approximately 9% lower estimated API cost per task for Astra versus Fable 5.1, and roughly 63% lower cost than Fable 5.1 when compared to GPT-5.6 Sol’s task cost. Astra’s published API pricing is $10 per million input tokens and $50 per million output tokens.

The framing is deliberate: “trained to complete tasks in fewer tokens with fewer retries.” If accurate, the token-efficiency claim matters as much as the raw benchmark score, since real-world agent costs scale with retries and conversation length, not just per-token price.

OpenAI also states that Astra “occupies the majority of the cost-efficiency frontier on professional work and coding evaluations, including Terminal-Bench 4.0 and Artificial Analysis Intelligence Index.” That is a positioning claim, not a benchmark result, and should be treated as such until independent analysis confirms it.

Where Astra Still Has No Primary Benchmark

GPT-6 Astra does not yet have a published SWE-bench Verified or tau-bench score. All available data for Astra is from benchmarks OpenAI either designed the evaluation for or released in its own launch materials: Terminal-Bench 4.0 and OfficeQA Pro. FrontierCode 1.1 (53.3%) and DeepSWE 1.1 (74.1%) appear in third-party comparison tables but are not from independent runs of SWE-bench Verified.

Until SWE-bench Verified scores appear, Astra’s position on this site’s ticker remains with a null composite score, imputed from the median of benchmarked models.

OfficeQA Pro

OpenAI also cites OfficeQA Pro in its launch materials — an internal benchmark that tests whether agents can find and analyze information across U.S. Treasury Bulletins, including financial tables, charts, and footnotes. GPT-6 Astra scores 69.9%, GPT-5.6 Sol scores 60.2%. The benchmark is not independently administered; it is OpenAI-designed and OpenAI-reported.