Training AI models with other AI models has become a popular goal for new labs, and now an Anthropic fellow has shown an early version of how it works in practice.
On Friday, the company published a paper titled “Automated Researchers Can Reliably Mitigate Alignment Failures”. It details how AI systems could reliably improve a model’s performance on a set of alignment benchmarks. When given 10 benchmarks for specific misaligned behaviors, the automated systems improved performance on every single one without degrading overall performance.
Led by Anthropic fellow Chen Yueh-Han, the system replicates much of the traditional approach to research. Each automated system searches the available literature, proposes a method, and trains the model using that method for 30 minutes. The benchmark increases gradually over several iterations. Effective methods are preserved while ineffective ones are discarded. This allows the system to operate quickly and at a great scale.
“Overall, these results provide early evidence that automated alignment post-training could become practical in the near term,” the paper reads.
The paper is a step toward recursive self-improvement, which many see as the next significant step in AI progress. If models can improve their own alignment training, it is plausible they could improve training practices more broadly. At that point, human AI researchers might soon become obsolete.
The paper is not shy about addressing this idea, explicitly comparing the Automated Alignment Researcher (AAR) to its human equivalent. “The best AAR method beats what experienced humans propose, on average within six hours,” the paper reads. “Human guided research directions do not lead to stronger performance.”
There is even a cost comparison, in case anyone was not convinced. “An AAR costs roughly $4 per hour in API inference against the $150 per hour we pay our human researchers.”
In fairness, the paper also points out a few limitations to this approach. The automated system only works insofar as the benchmarks reflect the actual alignment goals. Even then there is significant work to be done in establishing and maintaining those benchmarks — not to mention maintaining and expanding on the literature the automated researchers are drawn from.
What it means
For the people building these tools, the implication is clear. If software can outperform humans at finding ways to align models, the role of the human researcher shrinks to oversight. The cost difference between an automated system and a person is stark enough to change how labs allocate budget. The work now shifts to building better benchmarks and keeping the literature the systems read up to date.




