Claude Leads 26% of Anthropic R&D Tasks—but Human Oversight Still Defines “Leads”

On September 17, 2026, Anthropic published its internal AI-development measurement framework, which put Claude at the “leads” level for 26% of the company’s weighted AI research and development work as of August 2026. In this framework, leadership means completing most of a task end to end from a high-level prompt while a person continues to supervise.
AP’s September 18 account corroborated that about 90% of the measured work involved at least collaboration with Claude, while no measured portion had reached fully autonomous operation. The disclosure therefore shows broad and sometimes extensive use of Claude inside Anthropic’s model-development process, but not a system independently setting research priorities, validating its own conclusions or deciding what to deploy.
The percentages are company-reported internal measurements, not an independently audited assessment of Anthropic’s laboratory. Their meaning depends on the task basket, weighting method and boundary drawn between close collaboration and AI-led execution.
“Leads” is one rung below full autonomy
The automation index uses a six-step ladder running from work with no AI involvement to work performed without a human in the loop. In the middle, “assists” covers narrower support and “collaborates” covers substantial work completed under close human direction. “Leads” is the next rung: a person supplies the broad objective, but Claude handles most of the execution and works through complications without requiring continuous intervention.
Human authority remains part of that definition. Anthropic illustrates the distinction with a failed data pipeline: at the leadership level, Claude can inspect the failure, find its cause, write and test a repair, address additional problems and document the result. An engineer still reviews the work and decides whether the change enters production.
That example separates operational independence from institutional autonomy. Claude can carry a bounded assignment across multiple stages, but the framework does not credit it with independently deciding that the assignment should exist, accepting the risk of its result or authorizing deployment.
The denominator is a weighted basket of internal work
The headline share is not the percentage of all activity at Anthropic, nor a count of tickets, agent sessions or research decisions. The index begins with a catalogue of granular model-development tasks drawn from a staff sample and internal work records, then groups those tasks into a frozen hierarchy.
Categories receive weight according to the employee time devoted to them. A person’s weekly weight is divided evenly across the tasks attributed to that person, so work involving more staff time contributes more heavily to the final result. The published percentage is the weighted portion of that basket rated as AI-led.
Claude is also part of the measurement process. One model agent researches how a category of work is performed, and a separate Claude judge assigns an automation level; employees responsible for the relevant areas provide an independent comparison. This check can expose large disagreements, but it does not remove judgment from borderline cases where collaboration shades into leadership.
The frozen basket creates another limitation. It can show that work found in the baseline is becoming more automated, but it may not capture entirely new work that employees begin doing after older tasks move to Claude. A comparison with an earlier basket found no rise in novel task categories at the chosen level of analysis, although the methodology anticipates rebuilding and versioning the basket over time.
Collaboration and leadership are nested, not competing totals
The collaboration-or-higher share includes the AI-led share. It should not be read as one workload handled jointly by people and Claude alongside a separate workload handled by Claude alone.
At the collaboration level, a person remains closely involved and may need to resolve unexpected choices. At the leadership level, the person can step back during execution, but supervision and final responsibility remain. Full autonomy would require Claude to notice the need for work, scope it, perform it, test it and deploy the result without required human involvement—a threshold the measured basket did not reach.
This is also why the metric does not establish autonomous scientific discovery. It measures how work is carried out after tasks are identified and weighted, not who chooses Anthropic’s research agenda or whether Claude can independently verify that a research result is sound.
Thirty thousand concurrent agents measures scale, not control
Quartz’s coverage of the internal metrics records approximately 30,000 research and engineering agents operating simultaneously on Anthropic’s most-used internal platform in August 2026; it also notes that, among more than one billion agent decisions examined that month, the online monitor blocked 0.002%, or roughly one in 47,000.
The concurrent-agent figure is not a count of independent research projects, employees or autonomous systems. Agents can operate in parallel, communicate with one another and contribute to work that remains directed or reviewed by people. It describes operating scale, while the automation index describes how responsibility is divided within a weighted set of tasks.
The disclosed platform uses online monitoring before actions execute and offline monitoring afterward. Automated coverage across the activity stream is broader than human review: people examine blocked actions and the highest-priority cases that remain after additional filtering.
Oversight metrics show reach and speed, not detection quality
The proposed oversight framework separates three questions. Coverage asks how much agent activity passes through a monitor; review latency asks how long automated and human review take; and escalation rate tracks how often activity is blocked, redirected or flagged.
Those measures can reveal whether monitoring infrastructure expands as more work is delegated. They cannot, by themselves, show whether the monitors recognize every important failure. A low escalation rate could reflect rare misbehavior, weak detection or some combination of the two.
A further limitation is the short operating history: the monitoring setup has run for months rather than years, and the available evidence cannot establish that every consequential pattern is legible to human overseers. Independent evaluators, tests using known examples of misbehavior and repeated publication would provide stronger evidence than coverage figures alone.
For now, the defensible conclusion is narrower than the headline may suggest. Claude performs much of some measured R&D work from broad instructions, but people still define assignments, supervise outcomes and control deployment. Whether oversight quality is improving as quickly as task delegation remains unresolved.
Also read:
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.