Anthropic wants you to know Claude leads a quarter of its research, but “lead” doesn’t mean what you think

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane September 18, 2026 4 min read
Anthropic wants you to know Claude leads a quarter of its research, but “lead” doesn’t mean what you think

Anthropic reports that its AI, Claude, handles a quarter of the work on its next models. The company released figures showing Claude “leads” 26 percent of development tasks, a jump from under one percent in February. This data comes with significant caveats regarding scale, scoring methods, and the definition of the term “lead”.

In a company post, Anthropic outlined three specific measures. These include the extent to which AI handles model development alone, the ability to monitor AI agents, and the volume of computing power dedicated to safety. The firm backs each figure with internal operational data.

The release follows a call from Anthropic CEO Dario Amodei to slow development at the AI frontier in a coordinated way. The company argues the public needs more insight into how models are built. These metrics are intended to complement capability tests that measure what models can do.

Defining the role of the AI

The core of the report is an index sorting all development work onto a scale from Epoch AI. This scale runs from AL0, meaning no AI involvement, to AL5, meaning fully autonomous. As of August 2026, 26 percent of work sits at AL4, up from under one percent in February. More than 90 percent of tasks reach at least AL3. Claude never hits AL5.

Epoch AI labels AL4 as “AI leads,” and Anthropic used that phrase in its headline. Anyone interpreting “lead” as full autonomy is incorrect. That status belongs to AL5.

An example from Anthropic clarifies the AL4 position. An engineer hands Claude a bug report. The AI analyzes, fixes, and tests it without asking questions, but it cannot ship the code. A human reads the report and makes the final decision. The task and direction still come from the person. The difference from AL3, where the AI “collaborates,” is mainly that Claude no longer stalls when it runs into a snag.

Claude scored itself in this process. Agents gathered evidence from Slack and internal documents, and another Claude model assigned the levels. Anthropic admits this “judge” could make the same mistakes as the system it is checking.

Where “collaborates” ends and “leads” begins remains unclear even to Anthropic. A cross-check by the company reveals this. Employees were asked to judge how automated their own work area is. When two people rated the same area, they landed on the same level only about a third of the time. The Claude model that assigned the official scores matched human judgment 59 percent of the time. Almost always, in 97 percent of cases, human and model were at most one level apart. But one level is the jump from “collaborates” to “leads.” Whether a task counts toward the 26 percent is often a matter of interpretation.

The number also defines what the 26 percent actually measures. Anthropic listed how employees spent their work time in July and scored each activity, with tasks that consume more time counting for more. So Claude leads a quarter of the work as measured by the human hours it takes. That says nothing about how many decisions Claude makes or whether it has a say in the research direction.

The figure benefits the company in two ways. It serves as a performance claim for the product and as an argument for Amodei’s call to slow down. Anyone accusing AI firms of using safety warnings to shape regulation in their own favor will find fresh material in how “leads” is defined.

Monitoring 30,000 agents

On Anthropic’s most-used internal platform, about 30,000 agents run at the same time, according to the report. A real-time monitor checks every action and is meant to block things like copying model weights. Out of more than a billion decisions in August, it stopped 0.002 percent. A second monitor flags about 100,000 logs a week after the fact, and roughly 50 reach humans. Anthropic admits it has only run this monitoring for a few months and is not sure it catches every behaviour pattern.

During a sample week in July, about six percent of the compute for AI research went into safety work. Anthropic cares about this metric mainly because compute is the input that outsiders can check most easily. If the industry agreed to slow down, it would be a possible lever, for instance through voluntary commitments on the share going to safety research.

The low number does not necessarily mean little safety work. Anthropic says safety work is mostly about designing experiments, which costs researcher time but few chips, while a training run for a new model burns enormous capacity. Anthropic also did not count work that advances safety and capabilities equally as safety.

The company warns that the line to capability research is blurry, and every vendor is tempted to draw it generously. The burden of proof, it says, should sit with the developer.

What it means

For researchers, the shift to AL4 means the AI can finish a task end-to-end without human intervention, but a human must still approve the result before it leaves the system. The 26 percent figure highlights efficiency in execution rather than autonomy in decision making. It suggests the team spends less time on repetitive coding and more on architecture and oversight.

Scroll to Top