GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
GPT-56T 861 —
MUSE-SPK 835 -0.7%
GPT-56SC 828 -5.2%
QWEN-38X 824 —
CL-OP55X 822 —
GROK-46H 822 -5%
GPT-6A 820 —
GLM-5 784 -8.4%
CL-FAB5H 743 -5.6%
KIMI-K3X 742 -8.4%
CL-OP5H 720 -5.8%
CL-OP5X 709 -18%
CL-OP46H 698 -5.9%
CL-OP47H 690 -5.9%
GEM-38FH 677 +0.1%
GEM-37FH 657 -24%
GPT-56S 622 —
CL-OP47 582 -0.7%
GPT-55H 582 —
INKL 531 —
GEM-31P 513 —
GEM-3P 499 —
CL-OP46 496 -0.2%
CL-OP48 490 —
← Back to feed

Anthropic's New R&D Index: Claude Leads 26% of Internal AI Work, 90%+ at AI-Collaborates Level

Anthropic’s Institute published a measurement framework today for tracking how much of frontier AI development is being done by AI itself. The post introduces the Anthropic R&D Automation Index and provides a current snapshot of its numbers.

As of August 2026:

  • Claude is not operating fully autonomously for any measured subset of AI R&D work
  • Claude “leads” (AL4) 26% of Anthropic’s AI R&D — meaning it can complete most of a task end-to-end from a high-level prompt, with human supervision
  • More than 90% of AI R&D work sits at or above the “AI collaborates” threshold (AL3), meaning AI handles large chunks under close human direction

The automation level (AL) scale runs from AL0 (no AI involvement) to AL5 (fully autonomous, no human in the loop). AL4 is “AI leads”; AL3 is “AI collaborates.”

How the Index Was Built

Anthropic built the index by cataloguing every kind of AI R&D work done at the company, rating how automated each task currently is, and aggregating those ratings. The methodology samples staff weekly — for July 2026, roughly 20% of staff per department were sampled each week. A Claude research agent reviewed each sampled person’s week using Slack and internal documentation and listed the tasks they performed. That process generated approximately 15,000 granular model R&D tasks, which were then organized into a hierarchical tree.

The index is designed to be frozen at a baseline date. July 2026 is the reference basket. This means an increasing index number shows that the work humans were doing in July 2026 is being automated — it does not measure whether new categories of human work are emerging. Anthropic ran a parallel check using a January 2026 basket and found no rise in new task categories arriving monthly from February to July 2026.

The Self-Reference Problem

Anthropic used Claude to conduct the weekly audits of its own employees’ work. This creates a structural issue: the judge model could make the same kinds of errors as the model it is assessing. Anthropic acknowledges this and proposes third-party verification as a solution, specifically independent evaluators embedded inside the company with access comparable to internal risk assessment teams.

Why This Is Different From Earlier Announcements

In June 2026, Anthropic noted that Claude was writing more than 80% of its own production code. That figure was about one category of output — code. The R&D Automation Index covers all AI R&D work: experiments, evaluations, safety research, alignment work, infrastructure, and research operations. The 26% “AI leads” figure applies across that entire scope.

The new paper also introduces a public methodology that other frontier labs could replicate, and explicitly invites cross-lab comparison — subject to agreeing on a common task taxonomy and third-party verification.

Anthropic states it plans to report these numbers regularly and is proposing transparency obligations of this type as part of its Advanced AI Framework policy proposal.