GPT-6 Astra Posts 57.9% on Terminal-Bench 4.0, 2.1 Points Ahead of Fable 5.1
OpenAI on September 10 published a work-oriented launch post for GPT-6 Astra that includes the first Terminal-Bench 4.0 comparison table with Fable 5.1 and GPT-5.6 Sol in the same row. GPT-6 Astra scores 57.9%, Claude Fable 5.1 scores 55.8%, and GPT-5.6 Sol scores 37.3%.
What Terminal-Bench 4.0 Measures
Terminal-Bench 4.0 is a harder successor to the 2.x series. Cognition’s SWE-2, which scored 92.8% on Terminal-Bench 2.1, scored 27.3% on Terminal-Bench 4.0 when independently tested — a 65-point gap that exposed TB2.x as undersaturated at the frontier. TB4.0’s harder task distribution means the current cluster of frontier models sits in the 27–58% range rather than the 80–92% range TB2.x produced.
The current TB4.0 snapshot across published scores:
| Model | TB4.0 Score |
|---|---|
| GPT-6 Astra | 57.9% |
| Claude Fable 5.1 | 55.8% |
| GPT-5.6 Sol | 37.3% |
| Cognition SWE-2 | 27.3% |
The 2.1-point gap between Astra and Fable 5.1 is within the margin of error for most agent benchmarks, which typically carry ±2% confidence intervals. OpenAI does not include error bars in its launch post.
Cost Framing
OpenAI’s comparison highlights approximately 9% lower estimated API cost per task for Astra versus Fable 5.1, and roughly 63% lower cost than Fable 5.1 when compared to GPT-5.6 Sol’s task cost. Astra’s published API pricing is $10 per million input tokens and $50 per million output tokens.
The framing is deliberate: “trained to complete tasks in fewer tokens with fewer retries.” If accurate, the token-efficiency claim matters as much as the raw benchmark score, since real-world agent costs scale with retries and conversation length, not just per-token price.
OpenAI also states that Astra “occupies the majority of the cost-efficiency frontier on professional work and coding evaluations, including Terminal-Bench 4.0 and Artificial Analysis Intelligence Index.” That is a positioning claim, not a benchmark result, and should be treated as such until independent analysis confirms it.
Where Astra Still Has No Primary Benchmark
GPT-6 Astra does not yet have a published SWE-bench Verified or tau-bench score. All available data for Astra is from benchmarks OpenAI either designed the evaluation for or released in its own launch materials: Terminal-Bench 4.0 and OfficeQA Pro. FrontierCode 1.1 (53.3%) and DeepSWE 1.1 (74.1%) appear in third-party comparison tables but are not from independent runs of SWE-bench Verified.
Until SWE-bench Verified scores appear, Astra’s position on this site’s ticker remains with a null composite score, imputed from the median of benchmarked models.
OfficeQA Pro
OpenAI also cites OfficeQA Pro in its launch materials — an internal benchmark that tests whether agents can find and analyze information across U.S. Treasury Bulletins, including financial tables, charts, and footnotes. GPT-6 Astra scores 69.9%, GPT-5.6 Sol scores 60.2%. The benchmark is not independently administered; it is OpenAI-designed and OpenAI-reported.