GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Artificial Analysis Triples Terminal-Bench Weight in Engineering Index, Upgrades Six Domain Benchmarks

Artificial Analysis released Capability Indices v1.1 on September 14, updating its domain-specific AI benchmark suite across six professional verticals. The headline change: Terminal-Bench’s weight in the Engineering index triples from 5% to 15%, while GPQA Diamond reasoning exits entirely. The update shifts what it means to score well as an engineering AI.

Engineering Gets a Practical Tilt

The Engineering index previously weighted agentic terminal use at 5% and pure reasoning (GPQA Diamond and CritPoint) at 35% combined. Under v1.1, Terminal-Bench v4.0 replaces the older v2.1 and takes 15% of the index, with reasoning dropping to 30%. GPQA Diamond is removed entirely.

The practical implication: models that score well on abstract reasoning tests but struggle with actual terminal-based engineering tasks will rank lower. Terminal-Bench v4.0 is a harder version of the benchmark that measures autonomous execution of real engineering workflows in a bash environment.

AutomationBench-AA Enters Three Indexes

AutomationBench-AA, Artificial Analysis’s own agentic tool-use benchmark, was added at 10% weight to three indexes that previously had no agentic tool component:

  • Finance and Accounting: Adds AutomationBench-AA (Finance vertical) alongside AA-Briefcase and the GDP.pdf long-context test. Removes the τ³-Banking customer interaction benchmark.
  • Legal: Adds AutomationBench-AA (Operations vertical) with a new long-context component.
  • Strategy and Operations: Same additions as Finance, with τ³-Banking customer interaction also removed.

The removal of τ³-Banking from Finance and Strategy is notable. That benchmark measured agentic customer interaction capability; its removal suggests Artificial Analysis is shifting these indexes toward knowledge work and tool execution rather than conversational agent performance.

Healthcare Adds Long-Context Reasoning

The Healthcare and Medical index gains a new 15% long-context reasoning component via MLCR-AA while cutting non-hallucination weight from 15% to 10% and reasoning from 15% to 10%. Medical AI tasks often require synthesizing lengthy clinical documents; the MLCR-AA component directly targets that capability.

What v1.1 Rewards

Across all six indexes, the pattern is consistent: v1.1 shifts weight from passive knowledge evaluation toward active execution. The benchmarks that gained weight (Terminal-Bench, AutomationBench-AA, long-context reasoning) all require the model to do something with information rather than retrieve or classify it.

That matters for enterprise adoption decisions. When the benchmark score moves with actual task performance, it becomes a more useful signal for procurement. Artificial Analysis explicitly maps its Capability Indexes to O*NET occupational task distributions, so the weight changes track real-world job-function changes in what AI tools are being asked to do.

Full methodology: artificialanalysis.ai/methodology/capability-indices