Anthropic's research arm, the Anthropic Institute, published its first R&D Automation Index on Sept. 17: a measurement of how much of the company's own AI research and development work Claude now performs. As of August 2026, Claude operates at AL4 -- "leads" -- for 26% of measured R&D work, meaning it completes most such tasks end-to-end from a high-level prompt while a human supervises. That's up from under 1% in February. More than 90% of all measured work sits at or above AL3 ("collaborates," handling large chunks of a task under close direction), and Claude registered zero measured tasks at AL5 -- full autonomy, no human in the loop -- as of the same snapshot. Anthropic frames the release as a prototype methodology, not a finished scorecard -- the company's own stated goal is to give outside observers, including regulators and rival labs, a concrete way to talk about AI's role in AI development instead of arguing from anecdote. No other frontier lab -- not OpenAI, not Google DeepMind, not xAI -- publishes anything comparable about its own internal research process.
The methodology is unusually specific for a company self-report. Anthropic used the Automation Level (AL0-AL5) scale developed by Epoch AI, then built its own task taxonomy: in July, the company randomly sampled 20% of employees each week across departments involved in model R&D, and had a Claude research agent comb through Slack logs and internal documents to catalog what those employees actually did that week. Four weeks of sampling produced roughly 15,000 granular tasks, organized into a 542-node work-classification tree (378 leaf nodes) that stayed frozen for the duration, so the same yardstick applied to every measurement. The company also disclosed that roughly 30,000 agents are doing research and engineering work at any given moment on its most-used internal platform. Epoch AI, the independent research group that built the underlying AL0-AL5 scale, was not involved in scoring Anthropic's own tasks against it -- the classification pipeline that turned 15,000 sampled tasks into a single 26% figure was itself run by a Claude research agent, grading work partly performed by other Claude agents.
What Anthropic's own numbers cover -- and don't
- 26% · AL4 "leads" share, Aug. 2026
- Share of sampled R&D tasks Claude completes end-to-end under human supervision
Includes: Tasks from the 15,000-item, 4-week July sample, rated by a judge-model classifier against the frozen taxonomy.
Excludes: Any measure of code quality, research correctness, or output value -- this counts task completion under supervision, not whether the work was good. - 0.002% · Online-monitor block rate
- Share of agent actions blocked in real time (about 1 in 47,000)
Includes: Live blocks across the ~30,000 concurrently running internal research/engineering agents.
Excludes: Offline review, which runs separately and flags 1-2 transcripts per 1,000 for human follow-up -- a different, slower check on different failures. - 6% / 12% · Compute allocated to safety
- Share of total AI R&D compute, and of AI-driven R&D compute specifically, spent on safety work
Includes: Whatever Anthropic itself classifies as safety-directed compute spend.
Excludes: Independent verification of that classification -- the boundary between 'safety' and 'capabilities' compute is, by Anthropic's own admission, judgment-dependent.
The 30,000-concurrent-agents figure is the number likeliest to travel furthest out of context. It describes headcount on Anthropic's internal research and engineering agent platform at a single snapshot, not a claim about total autonomous decision-making capacity -- the same Index that discloses it also states plainly that none of that activity reached full autonomy (AL5) on any measured R&D task. Read together, the two figures describe a company running a very large number of closely supervised AI research assistants, not an unsupervised one. The Index also discloses how those agents are checked: online monitors block about 0.002% of agent decisions in real time, and a separate offline review flags 1-2 transcripts per 1,000 for human follow-up. Anthropic puts 6% of total AI R&D compute and 12% of AI-driven R&D compute specifically toward safety work -- the two figures the ledger above scopes by what counts as 'safety' in the company's own accounting.
The pacing essay and the automation index, dated
- Feb 2026 — Anthropic's baseline measurement: Claude "leads" under 1% of sampled AI R&D work.
- Jul 2026 — The Institute samples ~15,000 R&D tasks across four weeks to build the Index.
- Sept 12, 2026 — Amodei publishes "We Must Pace the Frontier," calling for industrywide deceleration; Altman, Musk and Hassabis endorse it the same day.
- Sept 17, 2026 — The Anthropic Institute publishes the R&D Automation Index: Claude now leads 26% of R&D work, up from under 1%.
- Sept 18, 2026 — Reuters reports Anthropic is weighing a new frontier model launch to counter GPT-6 Astra's enterprise momentum, ahead of a possible IPO.
Read next to each other, the essay and the Index measure genuinely different things -- one is about the pace of shipping more capable *models to the outside world*, the other about how much of Anthropic's *internal* research process Claude itself now performs. Anthropic could, in principle, hold its external release cadence steady while automating more of the research that happens before a release. But the two claims share a subject, arrived five days apart, and Forkast's own analysis of the pairing was blunt: "The entity most vocal about the risks of frontier AI is simultaneously the entity most aggressively automating its own development with that same technology," concluding that the gap "suggests that safety may be functioning more as institutional positioning than as a binding constraint on capability development." Anthropic itself has not publicly addressed the pairing; the Institute's post makes no reference to the pacing essay at all, and treats the Index as a standalone methodology release.
The Index's own fine print matters here too. Anthropic disclosed that its automation-level classifications -- assigned by a Claude judge model against the frozen taxonomy -- agreed with a human rater's classification of the same task 59% of the time. Two different human raters agreed with *each other* on the same task only 35% of the time. That's not a footnote: it means the underlying categories are genuinely fuzzy enough that trained people looking at the same unit of work often land on a different automation level than each other, let alone than the model. The 26% headline is real in the sense that Anthropic measured it consistently against its own frozen scale -- but it's a noisier number than a single clean percentage suggests, and no outside evaluator has yet re-run the same 15,000-task sample independently. Anthropic's own stated limitations go further: the 378-leaf task basket was frozen for the duration of the study specifically so the comparison would be consistent, which means it structurally cannot capture new categories of work that didn't exist when the taxonomy was built -- a real constraint on using this specific Index to track automation in fast-moving new research areas, even as it holds steady for tracking the categories it already defined.
- Anthropic says Claude now 'leads' 26% of its own AI R&D, up from under 1% in February.
- The Index sampled about 15,000 tasks across July, rating automation on a 6-level AL0-AL5 scale.
- Claude operated at full autonomy, AL5, on zero measured R&D tasks as of August, Anthropic says.
- Published 5 days after Amodei's industry slowdown essay; Reuters says a new model launch is weighed too.
- Caveat: two human raters agreed on the same task's automation level only 35% of the time.