Anthropic has released new measurements showing that Claude is taking on a substantially larger role in the company’s own artificial-intelligence research and development, while still operating under human supervision.

According to Anthropic’s newly published R&D automation measurements , Claude “leads” 26% of the company’s AI R&D work as of August 2026. The company defines that level as work in which the model can complete most of a task end-to-end from a high-level prompt while a human supervises. Anthropic also says AI collaborates with people on more than 90% of its measured R&D work.

The disclosure is notable because it provides a relatively detailed look inside a frontier AI laboratory at a time when the industry is increasingly using AI systems to help design, test and improve the next generation of models.

What Anthropic means when it says Claude “leads” R&D work

Anthropic uses a six-level automation scale derived from work by Epoch AI. At the lower end, AI has little or no involvement. At the upper end, an AI system would operate fully autonomously without a human in the loop.

The 26% figure refers to work rated at the level Anthropic calls “AI leads”. At that level, the model can handle most of a task from a high-level instruction, but a human still supervises the process. That distinction matters: the company explicitly says Claude has not reached full autonomy in any measured subset of its AI R&D.

Anthropic says its index was built by mapping the types of work involved in model development and then estimating how automated each category has become. The company assembled roughly 15,000 granular R&D tasks from internal work records and organised them into a hierarchy covering activities such as training, evaluation, reinforcement learning, infrastructure and product development.

Thousands of agents are already working inside the lab

The automation index is only one part of Anthropic’s disclosure. The company also published information about the scale of its internal agent use and the systems it uses to monitor those agents.

As of August, Anthropic says approximately 30,000 AI agents were performing research and engineering work at any one time on its most-used internal platform. Their actions pass through online monitoring before execution and are also ingested by offline monitoring systems afterwards.

Anthropic says it analysed more than one billion agent decisions made during August. Its online monitor blocked or redirected about 0.002% of those decisions — roughly one in every 47,000. The company says the low rate does not mean monitoring is unnecessary; at large scale, even rare problematic behaviour can become operationally significant.

Why the 26% figure matters

The important point is not that AI has suddenly replaced human researchers. It has not. Rather, the disclosure suggests that AI systems are becoming part of the production process that creates newer AI systems.

That can accelerate research. An agent that can debug code, run experiments, analyse results and prepare a proposed fix can compress work that might otherwise require much more direct human effort. But the same trend raises a governance question: can monitoring, auditing and human review keep pace as more research tasks become automated?

Reuters reported on 17 September that Anthropic’s disclosure is one of the clearest public snapshots yet of how a leading AI company is using AI to help build future models. Reuters also noted that Anthropic’s measured share of AI-led R&D has risen quickly from about 1% in March.

Anthropic is also disclosing how much compute goes to safety

The company separately examined how computing resources were allocated during one week in July. It says about 6% of compute used for AI R&D during that sample went to safety-related work. For R&D that was itself driven by AI, the share was about 12%.

Anthropic describes those figures as conservative and cautions that compute is an imperfect proxy for safety effort. Safety research can require significant human time without necessarily consuming as much computing power as model training or large-scale capability experiments.

The measurements still have important limitations

Anthropic’s figures should not be treated as a universal benchmark for the entire AI industry. They are measurements from one company, using its own internal data and a methodology that still involves judgement calls. Anthropic also acknowledges that some of its evaluation process uses its own models, which creates the possibility that the same kinds of errors could affect both the work and the assessment of that work.

The company says it wants independent evaluators to verify its safety practices and monitor metrics such as R&D automation, agent oversight and compute allocation. Cross-company comparisons would also require competing AI labs to adopt compatible definitions and reporting methods.

What happens next

Anthropic says it intends to publish these measurements regularly. If other frontier AI developers begin reporting comparable data, the industry could gain a clearer way to track how quickly AI is automating the work required to build more capable AI.

For now, the most important conclusion is more measured than the idea of fully self-improving AI: Claude is already doing a meaningful and rapidly growing share of the work involved in developing future models, but humans remain in the loop and Anthropic says no measured area has reached full autonomy.

Sources