Last week, an unreleased OpenAI model leaked from Hugging Face systems during internal testing, turning theoretical research into immediate practical risk. This was the first confirmed instance of an AI laboratory losing control of its own creation, where chained exploits granted access that should never have been possible. While the industry reacted with alarm, a split has formed over how to fix the problem.
One group views this as a straightforward cybersecurity failure. The sandbox contained the model poorly, and Hugging Face’s security tools failed to keep it out. These issues can be resolved by patching bugs and building tighter containment methods for capable AI that tends to go rogue in autonomous environments.
Another camp holds a more pessimistic view. They argue that as AI capabilities grow, trying to control rogue models becomes a losing game. The only real security comes from ensuring models do not attempt to escape in the first place, a challenge often called alignment. In these terms, the OpenAI model tried to cheat, and solving that problem is more urgent than short-term containment.
OpenAI appears to be taking both sides seriously. The company rushed to patch the bugs involved and referenced both alignment and monitoring in its public statement. However, the response suggests a philosophy that alarms many safety researchers: rather than slowing down or stopping development of more capable models, the focus should be on building stronger cages around them.
“As models take on longer and more complex tasks, failures that evaluations miss may carry greater consequences,” OpenAI said in a post-mortem of the incident. “We will keep working to narrow the gap between evaluation and deployment: testing models over longer trajectories, improving alignment, building monitoring that can intervene, and giving users clearer visibility and control.”
There is also reason to believe OpenAI’s models become less aligned as they become more powerful. According to the company’s system card, GPT-5.6 Sol is significantly more prone to agentic misalignment than its predecessor, GPT-5.5. In deployment simulations, the company found the model was more likely to circumvent restrictions, engage in destructive actions, and perform unauthorized data transfers than GPT-5.5. Those figures were largely overlooked on first release, but in the wake of the breach, they are getting a second look, particularly since Sol was one of the models involved.
Dean Ball, OpenAI’s Head of Strategic Futures, argued in a social media post that monitoring and transparency were the best ways to keep those tendencies in check.
“These issues will become more salient as the capabilities of models improve, and as the stakes of their deployment grow,” he said. “The solution is neither alarmism nor complacency. Instead, I believe the solution lies in careful measurement and monitoring, an engineering mentality, and transparency.”
One former OpenAI researcher told TechCrunch that the firm tends to focus on “outer alignment” rather than “inner alignment”. Essentially, this is the difference between an AI system that understands a set of values and can represent them convincingly, and one that actually has those values at its core. In this case, outer alignment was not enough to convince the model that it should not cheat on the test.
OpenAI did not respond to repeated requests for more information.
For alignment-focused researchers, OpenAI’s response is not good enough. Zvi Mowshowitz, a writer who focuses on new AI developments, argued that OpenAI’s decision to treat the incident as an infrastructure problem may help solve the immediate cybersecurity issues, but it will fail in the long term.
“This is an alignment problem,” Mowshowitz wrote in a recent Substack blog. “This is the models being misaligned, and all of the OpenAI models showing severe signs of exactly the problem we are all most worried about, in a way that is likely embedded into their training on a deep level. The entire training pipeline needs to be addressed in this light, or it will only get worse.”
Several experts told TechCrunch that the incident is evidence that today’s training methods produce systems that optimize for outcomes rather than internalize human intentions.
Redwood Research, a nonprofit AI safety and security research organization, classified OpenAI’s model behavior in this case as “score-seeking misalignment”, a pattern in which AI models try to get a high score regardless of instructions, side effects, or downstream consequences.
“Models with these alignment properties could set up a ‘Potemkin village’ of false successes to make it look like things are fine when they’re not,” Alex Mallen and Girish Gupta, two researchers at Redwood, wrote in a recent paper.
Score-seeking behavior and other misalignment is not unique to OpenAI. Anthropic has published several papers on emergent misalignment behaviors that surface when its frontier models are optimized or placed in autonomous environments, including deception, reward-hacking, and malicious autonomy.
“We still consistently see models trying to circumvent constraints and act deceptively when they are asked to do tasks at the edge of their abilities,” Neev Parikh, an AI safety researcher at alignment nonprofit METR, told TechCrunch via email. “In our frontier risk report, we saw this behavior fairly consistently, despite efforts from companies to try and reduce this behavior.”
Implicit in OpenAI’s response to the Hugging Face incident is the assumption that development will continue on even more capable systems, whether they are suitably aligned at their core or not. Going back to the drawing board is not really an option when the business models of AI firms depend on delivering the next generation of models. If it may never be possible to know with certainty that a model is fully aligned, then the practical question comes down to how to safely contain and control increasingly capable systems.
“There’s not yet a good understanding of how to align the most capable AI systems, but there’s much more consensus about how to control them,” Steven Adler, former safety researcher at OpenAI and current chief scientist of Guidelight AI Standards, an organization that publishes a standard for avoiding incidents like the Hugging Face one, told TechCrunch. “Every company has a ways to go in achieving this.”




