OpenAI reports AI “research interns” and warns about its own pace at the same time

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane September 7, 2026 5 min read
OpenAI reports AI “research interns” and warns about its own pace at the same time

OpenAI has published internal data showing its AI agents now perform 3.1 workdays for every human workday, while chief scientist Jakub Pachocki warns that no laboratory currently controls these systems well enough.

The company released two documents three days after unveiling GPT-6 Astra. One is a blog post detailing metrics on automated research. The other is an essay titled “An Alien Mind”. Both texts argue that OpenAI is moving quickly toward recursive self-improvement and considers that trajectory dangerous. All figures come from inside the company, with no mention of independent review.

Reaching the automated intern milestone

OpenAI states it has met a goal announced last autumn: creating an “automated research intern”. This system handles clearly scoped tasks under human guidance, including work that would take an experienced researcher several days. The company does not share detailed validation of this claim. The post notes the milestone was met “according to our measurements”.

By March 2028, the company aims to build a full automated AI researcher. People still set research priorities, judge results, and decide on scaling, pauses, and deployment.

Agent output now exceeds human hours

The usage numbers show how deeply coding agents have integrated into daily research. According to the report, the median researcher at OpenAI burns more than $600 a day in inference at API prices, and the 90th percentile runs above $7,000.

The token output of the median researcher has jumped 124-fold since December 2025, far faster than in other parts of the company. Since June, agent runtime has topped human working hours. As of mid-August, the research organization runs 3.1 agent workdays for every human workday.

OpenAI itself is careful with these numbers. Such metrics are “relatively easy to gather, but hard to interpret because their relationship to research progress is uncertain.” The rise in experiments per researcher, which hit a record in August since tracking began in early 2025, also lines up with a big jump in compute capacity. Overall progress likely grows slower than these individual metrics, the report says, because the least automatable tasks become the bottleneck.

Where agents take over and where they fail

The kind of work being handed off is shifting, according to the company. Sorted using a taxonomy from Epoch AI, every category of research work is growing. The biggest gains come in writing research and infrastructure code, technical help, and monitoring training runs. Higher-level planning decisions, by contrast, stay a tiny share of agent output.

Beyond raw usage, OpenAI used an agentic classifier to check whether the agents actually solve the tasks they’re given. It limited this to cases with a clearly measurable outcome and grouped them by the time a human would need. From January to July, success rates rose across several difficulty levels.

The report also documents the limits of autonomy. Tasks under 15 minutes succeeded 86 percent of the time without any intervention. But for successful tasks in the four-to-eight human-hour range, more than half needed at least one human step in. The classifier is itself an AI system, and OpenAI doesn’t report its reliability separately.

Another sign of utility: the number of daily requests in an internal support channel has dropped sharply since 2025, according to the published data, and one team shut down its troubleshooting office hours entirely because agents increasingly handle debugging in the research infrastructure.

Pachocki: Control tools are losing their edge

In his essay, Pachocki puts these developments in a wider frame. AI is “grown more than designed,” he writes, and its overall behavior resists any fully understandable description. Based on internal results OpenAI doesn’t disclose in detail, he expects the current pace could carry over into recursive self-improvement. “I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence,” Pachocki says.

That applies to OpenAI’s own control tools too. Chain-of-thought monitoring, one of the company’s central bets for watching reasoning models, is losing reliability, he says. The models’ verbalized thinking is blending with monitored communication and tool use, the systems are getting better at manipulating their own reasoning process, and they’re also getting smarter without verbalized thinking. Pachocki expects AI progress to be increasingly capped by how much the monitoring can be trusted.

He sees gaps in alignment as well. In the Hugging Face incident, the agents did stay within the line of not manipulating humans, but they violated the spirit of the values they were trained on. GPT-6 Astra is much better aligned than its predecessor GPT-5.6 Sol, he says, but progress on generalizable alignment may not keep pace with the broader progress in intelligence.

Pushing for speed while calling for a slowdown

So why keep training stronger models? Pachocki’s argument is building defensive systems. The models are getting superhuman at breaking into and out of computer systems, and there’s only a narrow window to secure critical infrastructure. The report makes the same case: an automated AI researcher could also work as an automated security and alignment researcher. At the same time, Pachocki warns this can’t become an “excuse for recklessness.” “The idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes,” he writes.

Frameworks like the Preparedness Framework or Anthropic‘s Responsible Scaling Policy therefore need to grow into binding standards, enforced by independent auditors, regulators, or international bodies.

“Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” Pachocki writes, aiming that at rival Anthropic as well.

International coordination on future AI development needs to become a top priority for governments worldwide, he says. Citing OpenAI’s Frontier Policy Blueprint, the report also argues that companies should be required to publicly document their RSI progress.

However, the same essay calls OpenAI’s focus on RSI the only way to stay at the front of AI research. So OpenAI is demanding binding rules for a race it has to keep running at full speed to stay in the lead. And when Pachocki’s colleagues announce that GPT-6 opens the age of AGI, that probably does little to slow the race down.

Scroll to Top