Claude Now Leads 26% of Anthropic’s AI R&D as Internal Agents Scale Up

· · Views: 2,121 · 3 min time to read

Anthropic says Claude now leads more than a quarter of the research and engineering work involved in building its next AI systems, offering a rare quantitative view of how quickly frontier laboratories are automating their own development processes.

Reuters reported that Claude now leads 26% of Anthropic’s measured AI research and development work, up dramatically from around 1% earlier this year. The company stresses that its models are not performing measured research fully autonomously and remain under human supervision.

Anthropic released the figures as part of a new effort to measure how much AI contributes to its own development.

More than 90% of research now involves AI collaboration

Anthropic created a scale ranging from AL0, where no AI is involved, to AL5, where AI operates fully autonomously without a person in the loop.

At AL3, AI collaborates on large pieces of work under close human direction. At AL4, AI “leads” by completing most of a task from a high-level prompt while a human supervises.

As of August, the company says Claude “leads” 26% of its AI R&D work while more than 90% is at or above the collaboration level. No measured category currently reaches fully autonomous AL5 work.

The measurement is based on roughly 15,000 granular R&D tasks collected from sampled employee work in July and organized into a hierarchy spanning areas such as pretraining, reinforcement learning, product development, evaluation infrastructure and serving incidents.

Anthropic says around 30,000 agents are active internally

The scale of internal agent deployment is equally striking.

Anthropic reported that approximately 30,000 AI agents were performing research and engineering work at any one time on its most-used internal platform in August.

Those agents operate with both real-time and after-the-fact monitoring.

Anthropic says online monitors evaluate actions before they occur and can block dangerous behavior, while offline systems review agent activity for slower-developing problems such as signs of misalignment.

Reuters reported that only around one in 47,000 monitored agent decisions was blocked for safety reasons, although that figure describes Anthropic’s own monitored deployment and should not be interpreted as a general failure rate for AI agents.

AI labs are becoming their own largest automation experiments

The numbers matter because frontier AI companies are not merely selling automation to customers.

They are increasingly using their own models to improve the systems that will replace them.

That creates a potential feedback loop: better models accelerate research, accelerated research produces better models, and each generation can take on a larger share of the work required to build the next.

Anthropic itself says the public needs greater visibility into how much AI is building the next version of itself as debate grows over whether frontier development should slow when oversight cannot keep pace.

The company also acknowledges limitations in its measurement approach, including the fact that Claude is being used to evaluate Claude’s own contribution to the work. Anthropic says third-party verification or evaluation by models from other developers could improve comparability.

For engineers and investors, the broader signal is substantial.

AI agents are beginning to move beyond isolated coding tasks into the research pipeline that produces the next generation of frontier models.

The question for AI labs may soon be less about whether they use agents internally and more about how quickly human supervision can scale when tens of thousands of agents are participating in the work at once.

Share
f 𝕏 in
Copied