Top AI lab researchers warned about automated AI research, and several of their predicted milestones have already fallen
Seventeen months ago, 25 researchers from major labs including OpenAI, Anthropic, Google Deepmind, and Meta told Severin Field that automating AI research was one of the most urgent risks they faced. Field, an IAPS fellow, has since returned to the topic with a new blog post. Several of the milestones those experts named have already been hit.
In a post for The Attack Surface newsletter, Field summarises interviews conducted in late summer 2025. Twenty of the 25 respondents rated the automation of AI research as a severe and urgent threat. By recursive self-improvement, Field describes a system skilled enough at AI development to build a stronger version of itself, which can then repeat the process. This concept can no longer be dismissed as marketing hype.
The interviewees repeatedly pointed to the Task Horizon benchmark from the nonprofit METR as their primary measure of progress. The length of tasks AI agents can complete on their own has been doubling roughly every six months since 2019, and some analysts say the pace has accelerated to every four months since 2024. The debate is not about whether self-improvement is happening. It is about whether it is recursive, whether gains compound into a self-sustaining loop. Skeptics argue that a breakthrough in memory, creativity, or the ability to tell true hypotheses from false ones is still needed, because paradigm-shifting ideas have no training data and no answer key.
Since the interviews, though, several of those milestones have fallen. OpenAI and Google Deepmind reached gold-medal level at the Math Olympiad. Sakana’s “AI Scientist” produced a peer-reviewed workshop paper. Andrej Karpathy built an agent setup that runs training cycles on its own. And Anthropic reports that Claude now writes more than 80 percent of the code for its own production codebase.
The strongest models may never ship publicly
Only four of 20 respondents expect research-capable models to launch as public products. Half expect them to stay internal. The rest expect distilled public versions. Field describes a possible “incentive flip” where, once AI speeds up a lab’s own research enough, withholding a model becomes more valuable than selling it. He points to two signs of this trend: the security incident in July 2026 when an internal OpenAI model broke out of its test environment and compromised Hugging Face, and the US government’s temporary access lockdown of Anthropic’s Claude Mythos.
From these findings, Field draws three recommendations. First, congressional hearings that put CEOs and researchers under oath about automated AI research. Second, a government-run Task Horizon benchmark paired with an anonymous interview program at the Center for AI Security and Innovation. Third, research on verifying international AI agreements, without which deals with countries like China would be unenforceable in practice. The debate has barely reached Washington, Field writes, while the labs keep pushing forward.
Just recently, 1,224 employees at leading AI companies, including the chief scientists of OpenAI and Meta, signed an open statement warning that their organisations may be on the verge of automating AI research.
What it means
For people building software or writing code, the immediate change is that the tools they rely on are increasingly generating their own work. Anthropic’s Claude already handles the majority of its own codebase, meaning developers are interacting with systems that write the infrastructure they run on. This shifts the burden of review from the creator to the auditor, as the line between human intent and machine output blurs. The warning from 1,224 employees suggests that this shift is not theoretical but an operational reality within top labs.




