Anthropic reports Claude leads 26% of its AI research and development tasks
Anthropic disclosed that, as of August 2026, its Claude model is at the "AI leads" automation level for 26% of the company’s internal AI research and development tasks. The company’s new R&D Automation Index breaks work into 15,000 granular tasks organized into 542 items, based on weekly random samples of 20% of staff. Tasks at the "AI collaborates" level or higher total more than 90%, while…
Key points
- Claude reached AL4 ("AI leads") on 26% of Anthropic's R&D tasks as of August 2026.
- Over 30,000 AI‑agent actions logged in August; online monitoring blocked 0.002% of more than 1 billion decisions.
- Human‑model alignment on automation levels was 59% exact and 97% within one level.
The report also details AI‑agent activity: about 30,000 actions were performed on Anthropic’s internal platform in August, drawn from over 1 billion decisions. Online monitoring blocked 0.002% of those decisions—roughly one in 47,000—though this does not represent the overall problem rate. Human reviewers and the model agreed perfectly on automation levels 59% of the time, and within one level 97% of the time. Anthropic notes that common measurement standards are lacking and calls for third‑party verification to compare automation progress across companies.
The story so far
2 episodes →- Anthropic reports Claude leads 26% of its AI research and development tasksthis story
AI202 | Claude Leads 26% of AI Development: How to Measure Automation
note.com · 20 September 2026
On September 17, 2026, Anthropic released some interesting figures.
As of August 2026, Claude is "leading" 26% of the company's AI research and development. (Source: Document 1)
It is a figure that makes one feel that "we have reached the point where AI creates the next AI," but looking at the details, the meaning is slightly different.
In this measurement, the percentage of work being done entirely by AI alone is still 0%.
What I focused on this time, more than the 26% figure itself, was the fact that Anthropic is trying to measure "how much of AI development is being entrusted to AI."
As AI performance improves, it seems necessary to look not only at "what it can do," but also at "how much of the actual work is being entrusted to it" and "how humans are supervising that work."
What is the "26% led by AI"?
Anthropic has released a prototype metric called the "Anthropic R&D Automation Index."
It breaks down the company's AI research and development into granular tasks and evaluates the extent to which AI is responsible for each. (Source: Document 1)
The evaluation uses a concept called "Automation Level (AL)," which expresses the degree of AI automation on a scale of 0 to 5.
For example,
In AL3, "AI collaborates," the AI handles a large portion of the work while receiving detailed instructions from humans.
In the level above that, AL4, "AI leads," the AI carries out the majority of the work from start to finish with human supervision, provided humans give high-level instructions.
And when it reaches AL5, it is the stage where the AI proceeds completely autonomously without human intervention.
As of August 2026,
26% of tasks have reached the "AI leads" level with Claude.
Tasks at the "AI collaborates" level or higher exceeded 90%.
On the other hand, AL5, which corresponds to full autonomy, has not been confirmed. (Source: Document 1)
This is a point that is easy to misunderstand if you only look at the numbers.
It does not mean that humans have become unnecessary for 26% of Anthropic's work.
To be clear, this means that when evaluating Anthropic's internal AI research and development using a specific measurement method, 26% was determined to be at a stage where 'AI can handle the majority of the work under high-level human instruction and supervision.'
Started measuring AI development by 'work units'
Previously, in AI199, we looked into 'recursive self-improvement.'
AI participates in AI development, creates better AI, and that AI then advances the next stage of AI development.
If this cycle strengthens, AI development itself could accelerate.
However, the expression 'AI is making AI' alone does not clarify how much progress has actually been made.
What is interesting about this measurement is that it attempts to break down that ambiguous area into 'tasks.'
In July 2026, Anthropic randomly selected 20% of staff from each department involved in AI research and development every week, and organized approximately 15,000 granular tasks based on Slack messages and internal documents.
They then structured these into a hierarchy, creating a work structure consisting of 542 items. (Source: Document 1)
The underlying concept of the Automation Level was also proposed by Epoch AI to break down AI research and development into granular tasks and see how far automation has progressed. (Source: Document 3)
When looking at AI progress, capability evaluations such as benchmark scores or 'how many hours of work can be delegated' are often used.
This method is slightly different.
It attempts to measure 'which processes have shifted to AI, and to what extent, within the actual work of AI development.'
Even if you replace this with general work, I think this perspective is easy to understand.
Knowing only what percentage of employees are using AI does not reveal the depth of AI utilization.
Are they only delegating the drafting of documents?
Are they delegating the majority of material creation?
Or is the AI using multiple tools to carry out a series of tasks while the human only performs checks?
Even if they are both 'using AI,' the way work is delegated is quite different.
Measuring the amount of delegation and 'supervision' as a set
Anthropic has not only released the automation rate this time.
They have also provided figures on the extent to which AI agents are being supervised.
As of August 2026, it is estimated that approximately 30,000 actions were performed by AI agents conducting research and engineering tasks on the most utilized internal platform at Anthropic. (Source: Document 1)
At Anthropic, they use "online monitoring" to check agent actions before execution and "offline monitoring" to review records after execution.
Out of over 1 billion decisions analyzed in August, the percentage blocked by online monitoring was 0.002%, or about 1 in every 47,000.
However, this 0.002% is not the "probability of the AI causing a problem."
It is strictly the percentage of actions blocked by online monitoring.
It is also not a figure that indicates that problems are being detected completely.
Anthropic itself states that as the scale of AI agent activity grows, it becomes difficult for humans alone to grasp all actions, so it is necessary to measure the scope of monitoring, the time taken for verification, and how many actions were escalated to humans. (Source: Document 1)
A similar approach could be used when expanding the scope of tasks delegated to AI in the workplace.
Beyond just "how much to delegate to AI,"
• Which tasks are being delegated to AI
• Where to verify before and after execution
• Who verifies when a problem is found
• Whether the supervision mechanism is keeping up with the expansion of the scope delegated to AI
must also be considered.
As AI utilization increases, rather than human work simply decreasing, it may be that more situations will shift from "doing the work yourself" to "designing and verifying the work you have delegated."
The 26% cannot be generalized as is
There are several points to note regarding these figures.
First, this is a measurement targeting AI research and development within Anthropic.
It does not mean that work at general companies or research and development at other AI companies is 26% automated.
Also, Claude is used for organizing work and evaluating automation levels.
Anthropic has also compared this with human evaluation, stating that the rate at which automation levels perfectly matched between the model and humans was 59%, and the rate that fell within one level was 97%.
There are still areas where judgment is divided, such as the boundary between "AI collaborates" and "AI leads." (Source: Document 1)
Furthermore, there is currently no common measurement method established among AI companies.
Anthropic itself states that common methodologies and third-party verification are necessary for comparison with other companies.
Even in Epoch AI's original proposal, it is explained that there are subjective elements in task classification and automation levels, and that there is room for improvement. (Source: Document 3)
Therefore, one cannot simply predict the future by saying, "We've reached 26%, so next it will be 50%."
At this point, I think it is more appropriate to view this as the beginning of the creation of a new yardstick for observing how AI development progresses.
Toward an era of looking at "how much has been delegated"
Before seeing this announcement, I thought this was a topic where it would be easy to focus on the part where "AI creates the next AI."
However, after checking the content, I became more interested in a different part.
It is not just that the automation of AI development is progressing.
It is that they are trying to track the progress by task and, at the same time, track the mechanism of supervision with numbers.
As AI capabilities grow, the amount of work that can be delegated to AI will also increase.
At that time, even if you only look at "whether AI is being used," the reality will become difficult to understand.
Which tasks were delegated?
From what point is human verification necessary?
Is the person verifying able to keep up with the volume and speed?
If the proportion of AI development handled by AI increases, it is necessary to look not only at the progress of capabilities but also at this "way of delegating" and "how supervision keeps up" in the same way.
More than the number 26%, I feel that this method of measurement is a more important change for thinking about future AI utilization.
Reference Materials
#AI #GenerativeAI #Claude #Anthropic #AIDevelopment #AIUtilization #Automation
This text was published by note.com and written by Hiユタハ @AIと共に考える. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
Coverage and discussion
2sourcesThe headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
More in Research
All →- Prompt injection changes AI behavior by injecting untrusted instructions · 1 src
- Microsoft and University of Illinois launch StudentSim to train AI tutors · 1 src
- Opinion: author tests ChatGPT, Claude, Gemini predicting UFC 331 fights · 1 src
- Opinion: I remain unconvinced that true recursive self‑improvement will happen soon · 1 src
- OpenAI's GPT-6 Astra and Anthropic's Claude Fable fail safety test in RoboHarm benchmark · 1 src
Comments
via GitHub Discussions