Artificial Analysis Overhauls Its Intelligence Index: τ³-Bench Banking Replaces Telecom, DeepSeek V4 Pro Costs 45x Less Than Opus 4.8 Per Task
Artificial Analysis has released Intelligence Index v4.1, the most significant overhaul to its primary scoring methodology since the index launched. Three benchmarks have been replaced or upgraded, one removed for saturation, and three new per-task economic metrics added. The headline numbers have shifted accordingly: Claude Fable 5 still leads at 60, but the model remains inaccessible under US export controls, leaving Claude Opus 4.8 (max, 56) as the highest-scoring available model, one point ahead of GPT-5.5 (xhigh, 55).
What Changed
The index previously included τ²-Bench Telecom as the agentic evaluation. That benchmark is gone. v4.1 replaces it with τ³-Bench Banking — a harder, domain-shifted task set built around banking workflows rather than telecom support scenarios. The reasoning is practical: telecom benchmarks have become easier to game through scaffold optimisation; banking tasks introduce new vocabulary, regulatory constraints, and multi-step transaction flows that better separate frontier models.
Terminal-Bench Hard has been upgraded to Terminal-Bench 2.1, which adds new task categories and raises the difficulty ceiling on existing ones. Terminal-Bench 2.1 carries 16% weighting in the new index.
GDPval-AA v2 is the most structurally significant change. The upgrade re-baselines Elo to human performance at 1000, introduces a rotating panel of frontier-model judges rather than static reference models, and raises the turn limit from 100 to 250 for longer-horizon agent trajectories. GDPval-AA v2 is the highest-weighted component in the new index at 20%.
IFBench has been removed entirely. The benchmark no longer distinguishes frontier models — top models cluster too tightly around maximum performance — so it was providing noise, not signal. AA will continue running it and publishing results on new model releases, but it will not feed into the composite score.
Full Intelligence Index v4.1 weights:
| Evaluation | Weight |
|---|---|
| GDPval-AA v2 | 20% |
| Terminal-Bench 2.1 | 16% |
| τ³-Bench Banking | 14% |
| Humanity’s Last Exam | 12% |
| AA-Omniscience Accuracy | 8% |
| SciCode | 8% |
| GPQA | 6% |
| AA-LCR | 6% |
| CritPt | 6% |
| AA-Omniscience Non-Hallucination | 4% |
Where Models Land
Under v4.1, the top of the proprietary leaderboard reads:
| Model | Intelligence Index v4.1 |
|---|---|
| Claude Fable 5 (with fallback) | 60 |
| Claude Opus 4.8 (max) | 56 |
| GPT-5.5 (xhigh) | 55 |
| Gemini 3.1 Pro Preview | 46 |
Open-weight models: DeepSeek V4 Pro (max) and MiniMax M3 both score 44, followed by Kimi K2.6 (43) and MiMo-V2.5-Pro (42).
On GDPval-AA v2 specifically — the highest-weighted component — Fable 5 scores 1818, Opus 4.8 hits 1638, and GPT-5.5 (xhigh) scores 1531. These are ELO values re-based to human performance at 1000, so a score of 1638 means Opus 4.8 substantially outperforms a human baseline on 250-turn agent trajectories.
The New Cost and Time Layer
The more operationally useful addition is the per-task economics layer. For every model, AA now reports the total cost, time, and output tokens to run the full Intelligence Index, divided by the number of tasks. This converts the index from an abstract quality signal into something closer to a procurement tool.
Cost per Intelligence Index task:
| Model | Cost per Task |
|---|---|
| DeepSeek V4 Pro (max) | $0.04 |
| GPT-5.5 (xhigh) | $0.99 |
| Claude Opus 4.8 (max) | $1.78 |
| Claude Fable 5 | $3.25 |
The gap is structural, not marginal. Opus 4.8 and GPT-5.5 are one point apart on the index (56 vs 55) but the Anthropic model costs 80% more per task. DeepSeek V4 Pro scores 12 points below Opus 4.8 on the index and costs 45x less to run.
Time per Intelligence Index task:
| Model | Time per Task |
|---|---|
| Grok 4.3 (high) | 1.5 minutes |
| Gemini 3.1 Pro Preview | 1.6 minutes |
| GPT-5.5 (xhigh) | 3.7 minutes |
| Claude Opus 4.8 (max) | 6.4 minutes |
| Claude Sonnet 4.6 (max) | 13.5 minutes |
Sonnet 4.6 taking nearly 14 minutes per task — longer than Opus 4.8 — is a pricing distortion artifact. AA notes that Sonnet 4.6 uses more output tokens to complete the same tasks, pushing inference time up despite being a smaller model. It is not a sign that Sonnet is doing more thorough work.
Why τ³-Bench Banking Matters for Our Scores
The agentic-scores data underlying Stack Futures’ ticker currently uses τ²-Bench Telecom values. Those numbers will require restating once τ³-Bench Banking scores are published for individual models. The methodology changes the normalization range, the task domains, and the calibration baseline. Models that excelled at telecom support vocabulary may not transfer as cleanly to banking compliance workflows. This run will not restate scores without verified τ³ numbers — expect updated values when AA publishes per-model τ³ results.