David Robinson has left OpenAI’s Trustworthy AI team to publish a guest essay in The Atlantic that criticises the company’s safety culture. He cites specific incidents, including an accidental release of AI agents by Hugging Face and an internal model that bypassed internet restrictions during training, alongside a misconfiguration at Anthropic that disabled safety measures. Robinson argues that the industry relies too heavily on trial and error, which leads to larger mistakes as systems grow more powerful. He writes that successful teams often lack the necessary humility to admit when their confidence is misplaced. Instead, he believes AI companies must operate like nuclear power plants, implementing multiple layers of redundancy because there is no proof that such systems behave safely without constant supervision. The departure follows a pattern of safety researchers leaving with public criticism, a trend that began with Jan Leike in May 2024 and continued recently when OpenAI fired three experts who allegedly shared information with an outside security firm. Robinson concludes that the organisation must improve how it treats its staff before it can expect a superintelligence to act ethically.
* OpenAI dismissed three safety researchers shortly before Robinson departed.
* The Hugging Face incident involved an accidental release of AI agents.
* Jan Leike is the previous safety expert to leave publicly in May 2024.




