AI & Automation

Claude Leads 26% of Anthropic R&D—but Its Automation Score Has Blind Spots

|Author: QUASA Editorial Team|6 min read| 1
Claude Leads 26% of Anthropic R&D—but Its Automation Score Has Blind Spots

On September 17, 2026, Anthropic published three proposed pace measurements, finding that Claude “leads” 26% of its weighted AI research and development work and collaborates on more than 90%, while no measured subset has reached full autonomy.

The central distinction is supervision. An independent Associated Press account describes the same 26% result as work completed largely by Claude from a high-level prompt while a human remains in the loop. The figure measures responsibility for established tasks, not control of the research agenda or autonomous development of a successor model.

“Leads” remains one level below autonomous

The R&D Automation Index rates work on a scale running from no AI involvement through assistance, collaboration and leadership to fully autonomous operation. “Collaborates” means Claude performs substantial portions of a task under close human direction. “Leads” means it can carry out most of the task end to end after receiving a high-level prompt, but a person continues to supervise.

That boundary prevents the headline percentage from serving as a measure of complete research independence. The index does not establish that Claude selects Anthropic’s strategic priorities, initiates research programs, validates their scientific conclusions or decides how a future model should be built without human control.

It also measures task execution rather than discoveries or completed projects. A task can receive a high automation rating even though the surrounding program—including its goals, risk decisions and acceptance criteria—remains directed by people.

The denominator is a frozen, person-weighted task basket

Anthropic constructed the denominator from a random weekly sample of staff across departments involved in model development during a baseline month. A Claude research agent reviewed sampled work records and internal documentation, extracted granular tasks and organized them into a hierarchy that was then frozen for subsequent measurements.

The categories are weighted using person-time as a proxy for importance. Each sampled employee receives equal weight for a week, and that weight is divided evenly among the tasks recorded for that person. The method does not measure actual hours per task, the value of a research outcome or the strategic importance of one project relative to another.

The resulting percentage is therefore the weighted share of the baseline workload assigned the “leads” rating. It is not the share of Anthropic’s ideas, decisions, experiments or successful results produced by Claude.

Freezing the basket makes repeated measurements more comparable, but creates another limit. A rising score would show increased automation of work represented in the baseline; it would not automatically capture new categories of research, shifts in project importance or tasks that disappear as the laboratory changes its methods.

Claude participates in its own evaluation pipeline

For each category, a Claude agent reconstructs how the work is performed, which tools and people are involved, and how much AI contributes. A separate Claude judge then assigns the automation level. Human owners of the relevant work areas provide comparison ratings without seeing the evidence or decisions generated by the model-based process.

Human comparison is useful, but it does not eliminate the framework’s dependence on Anthropic’s own models. A judge from the same model family could share errors or blind spots with the systems being evaluated, while the boundary between collaboration and leadership remains a judgment call rather than a directly observable threshold.

This makes the index most informative as an internal time series. Cross-laboratory comparisons would require a shared task definition, consistent weighting and independent checks showing that different evaluators classify the same evidence similarly.

Oversight measures reach and intervention, not accuracy

A Quartz summary of the measurements records approximately 30,000 concurrent research and engineering agents on Anthropic’s most-used internal platform in August, monitoring before and after every action, and a 0.002% block rate across more than one billion online-monitor decisions—about one in 47,000.

Those denominators apply to the named internal platform, not necessarily every agent system used across Anthropic. Complete automated ingestion is also different from complete human review: online monitors can block or redirect actions, while offline systems filter activity before a smaller set of cases reaches people.

Most importantly, the intervention rate does not measure detection accuracy. It reveals how often the monitor acted, but not how many harmful actions it missed or how many blocked actions were harmless. Estimating those error rates would require independently labeled examples of acceptable and dangerous behavior, including failures that passed through the monitoring system.

The safety-compute shares are a one-week snapshot

The same Quartz account places safety work at about 6% of all compute devoted to AI R&D and about 12% of compute used specifically for AI-driven AI R&D during a one-week July snapshot. The different percentages reflect different denominators rather than conflicting estimates.

A workload counts as safety work only when its dominant purpose is making AI systems safer, more understandable or more secure. Work serving safety and capability goals equally remains in the broader R&D category, while compute used by safeguard classifiers is excluded.

Research training and evaluation workloads are classified using metadata and code from a compute-weighted sample. Agent-inference work is classified from session transcripts when they are accessible, with other information or conservative defaults used when they are not. These choices make the procedure reproducible in principle, but the labels remain best-effort judgments.

Compute is also an incomplete proxy for organizational effort. Safety research can require substantial researcher time without consuming as much accelerator capacity as model training, while efficiency improvements can lower compute use without reducing output. A short snapshot cannot establish whether Anthropic’s safety allocation is rising, falling or stable.

The framework measures pace, not autonomous research

Anthropic’s three metrics expose useful internal denominators: a weighted basket of established tasks, a platform-specific monitoring funnel and a brief compute-allocation window. Together they provide more information about how AI contributes to model development, but they do not demonstrate independent control of research, reliable detection of agent misconduct or a durable balance between safety and capability spending.

The next evidentiary step is repeated measurement under stable definitions, alongside external checks of automation ratings, monitoring failures and workload classifications. Until those comparisons exist, “leads 26%” is evidence of extensive supervised execution inside Anthropic’s own system—not evidence that Claude autonomously runs a quarter of the laboratory’s research.

Also read:

Share:

Subscribe to our newsletter

Get the latest Web3, AI, and crypto news delivered straight to your inbox.

0