Artificial Analysis Triples Terminal-Bench Weight in Engineering Index, Upgrades Six Domain Benchmarks
Artificial Analysis released Capability Indices v1.1 on September 14, updating its domain-specific AI benchmark suite across six professional verticals. The headline change: Terminal-Bench’s weight in the Engineering index triples from 5% to 15%, while GPQA Diamond reasoning exits entirely. The update shifts what it means to score well as an engineering AI.
Engineering Gets a Practical Tilt
The Engineering index previously weighted agentic terminal use at 5% and pure reasoning (GPQA Diamond and CritPoint) at 35% combined. Under v1.1, Terminal-Bench v4.0 replaces the older v2.1 and takes 15% of the index, with reasoning dropping to 30%. GPQA Diamond is removed entirely.
The practical implication: models that score well on abstract reasoning tests but struggle with actual terminal-based engineering tasks will rank lower. Terminal-Bench v4.0 is a harder version of the benchmark that measures autonomous execution of real engineering workflows in a bash environment.
AutomationBench-AA Enters Three Indexes
AutomationBench-AA, Artificial Analysis’s own agentic tool-use benchmark, was added at 10% weight to three indexes that previously had no agentic tool component:
- Finance and Accounting: Adds AutomationBench-AA (Finance vertical) alongside AA-Briefcase and the GDP.pdf long-context test. Removes the τ³-Banking customer interaction benchmark.
- Legal: Adds AutomationBench-AA (Operations vertical) with a new long-context component.
- Strategy and Operations: Same additions as Finance, with τ³-Banking customer interaction also removed.
The removal of τ³-Banking from Finance and Strategy is notable. That benchmark measured agentic customer interaction capability; its removal suggests Artificial Analysis is shifting these indexes toward knowledge work and tool execution rather than conversational agent performance.
Healthcare Adds Long-Context Reasoning
The Healthcare and Medical index gains a new 15% long-context reasoning component via MLCR-AA while cutting non-hallucination weight from 15% to 10% and reasoning from 15% to 10%. Medical AI tasks often require synthesizing lengthy clinical documents; the MLCR-AA component directly targets that capability.
What v1.1 Rewards
Across all six indexes, the pattern is consistent: v1.1 shifts weight from passive knowledge evaluation toward active execution. The benchmarks that gained weight (Terminal-Bench, AutomationBench-AA, long-context reasoning) all require the model to do something with information rather than retrieve or classify it.
That matters for enterprise adoption decisions. When the benchmark score moves with actual task performance, it becomes a more useful signal for procurement. Artificial Analysis explicitly maps its Capability Indexes to O*NET occupational task distributions, so the weight changes track real-world job-function changes in what AI tools are being asked to do.