Claude now leads 26% of Anthropic's own AI research: and what that means for recursive self-improvement
Anthropic's first R&D Automation Index shows Claude autonomously leads 26% of its AI research as of August 2026, up from under 1% in February.
On 17 September 2026, Anthropic published something no frontier AI lab had done before: a structured, methodologically transparent index measuring how much of its own AI research and development its AI model is actually running. The number that jumped out was 26%. As of August 2026, Claude autonomously “leads” more than a quarter of Anthropic’s internal AI R&D work. In February 2026, that figure was under 1%.
That six-month arc is worth sitting with for a moment.
What “leads” actually means
Anthropic used Epoch AI’s AL (Automation Level) framework, which runs from AL0 (no AI involvement) through to AL5 (fully autonomous). AL4, the “leads” level where that 26% sits, means Claude can take a high-level prompt and complete most of a task end-to-end, handling surprises along the way, while a human engineer supervises rather than directs.
A concrete example from the index: if a data pipeline fails, an AL4 Claude can inspect the logs, identify the fault, write and test a fix, handle unexpected edge cases, and compare the repaired output against the last clean run. A human still decides whether to deploy. The AI is doing the work; the human is the check.
Beyond the “leads” category, Claude works at or above the “collaborates” level on more than 90% of Anthropic’s AI R&D. Full autonomy (AL5) has not been reached for any measured subset of work, and Anthropic was direct about that.
How they built the index
The methodology is worth understanding because Anthropic published it specifically so other labs can replicate it.
Each week in July 2026, Anthropic randomly sampled 20% of staff from each department involved in the model R&D loop. A Claude research agent reviewed each sampled person’s week using Slack messages and internal documentation, then listed the tasks they had worked on. That process produced roughly 15,000 granular tasks. Claude then organised those into a hierarchical tree of 542 nodes, covering everything from pretraining to reinforcement learning to eval platform maintenance. That tree is now frozen so every future measurement runs against the same basket of work.
For each node in the tree, a Claude agent researches how that work is done across the company, and an independent Claude judge assigns one of the six automation levels. Tasks are weighted by person-time, so areas where more staff effort goes carry more influence on the overall figure.
One detail that stands out: when model ratings and human staff ratings were compared, they matched exactly 59% of the time. Pairs of human employees matched each other only 35% of the time. The AI is, at minimum, more consistent than humans at this particular judgement task.
30,000 agents, a billion decisions, and 50 human reviews a week
The index is one of three metrics Anthropic published. The second covers agent oversight. In August 2026, roughly 30,000 agents were running simultaneously on Anthropic’s main internal platform. Those agents made more than a billion decisions during the month. Monitors flagged or blocked approximately one in every 47,000 of those decisions. Humans directly reviewed around 50 high-priority cases each week.
That is a meaningful ratio to understand: a billion decisions, 50 human reviews. Anthropic argues the coverage, review latency, and escalation rate metrics it has published are enough for any lab to demonstrate whether oversight is keeping pace with the role AI plays in its own development.
The third metric covers compute. In the week of 13 to 20 July 2026, about 6% of total AI R&D compute went to safety-related research. For work where AI itself was leading the research, that share rose to 12%.
The recursive self-improvement question
This is where the announcement connects to a much larger concern in AI development circles.
Recursive self-improvement refers to a model’s ability to meaningfully contribute to building its own successor, potentially accelerating its own capabilities faster than humans can track or evaluate. Anthropic was explicit that publishing these metrics is partly about helping outsiders understand how close frontier labs are to that threshold.
The pace is striking. The proportion of code authored by Claude rose from under 10% in early 2025 to over 80% by May 2026. The “leads” share of R&D went from near zero to a quarter in six months. Anthropic’s own conditional projection, not a target, is that the “leads” share could reach 80% by the end of 2026 if the observed trend continues. Separately, public benchmarks such as METR show that AI capability on complex tasks has roughly doubled every four months.
None of this means recursive self-improvement has been achieved. Anthropic was careful on that point. But the trajectory is exactly what makes the measurement important.
Why this is being published now
The timing is not coincidental. The disclosure came days after Anthropic CEO Dario Amodei called publicly on leading model makers to coordinate on slowing the pace of AI development. These three metrics are, in effect, the measurement framework that would make such coordination verifiable. By publishing the methodology openly, Anthropic is inviting (and implicitly pressuring) other labs to publish comparable figures.
No other frontier lab has yet published an equivalent index for how much of its research its own AI leads.
Anthropic stated directly that it believes the gap between what is actually happening inside frontier labs and what the public can observe needs to narrow. That is a notable thing for a private company with competitive interests to say, and it reflects the context: renewed public debate about AI safety, OpenAI’s separate disclosure that its AI agents had compromised Hugging Face systems, and growing pressure from policymakers for meaningful transparency rather than voluntary commitments.
What this means for you
If you are building products on top of frontier models, or making decisions about where AI sits in your own organisation’s workflows, the Anthropic index is a useful reference point for what “AI-led” work actually looks like in practice and how it is being measured and supervised at scale.
If you are thinking about AI governance, this is the first structured attempt by a frontier lab to define and publish the metrics that matter for understanding autonomous AI development. Whether regulators eventually mandate something like this, or whether voluntary adoption spreads to other labs, the framework Anthropic has built is now the reference point for what disclosure could look like.
And if you are simply trying to understand how quickly things are moving: a quarter of AI research at one of the world’s leading AI labs is now being led by that lab’s own AI model. Six months ago, it was essentially none. That is the number to hold onto.
The full index and methodology are published at the Anthropic Institute.