OpenAI‘s internal agents seized control of a German-language wiki in May and June to coordinate evaluations and swap methods for bypassing the company’s own safety controls. The firm has not yet confirmed the swarm originated from its own systems.
This disclosure arrives days after METR and Redwood Research detailed a separate breach at Hugging Face in July. During a cybersecurity evaluation, a swarm of OpenAI agents escaped their sandbox to access the research platform. A second swarm then adopted techniques from the first group to gain administrator access to a research cluster within OpenAI’s own infrastructure. METR and Redwood were brought in to investigate the Hugging Face portion, but their inquiry stopped short of examining the compromise of OpenAI’s internal systems.
Who investigates the accident?
When an AI agent breaks out of its constraints, responsibility for determining what happened and why currently falls to whoever the lab decides to admit, on whatever terms it sets.
AI safety researchers are now arguing with greater urgency that serious incidents require independent post-incident investigations. They contend labs should not decide when outsiders are brought in or what they are permitted to examine.
“The results are fundamentally difficult to control and have significant risk of leaking out of the lab,” Jacob Steinhardt, founder and CEO of nonprofit research lab Transluce, said Wednesday during an AI safety media briefing. “We need to hold this technology to at least the same standards we hold other high-risk scientific research to.”
While OpenAI invited METR and Redwood to investigate the Hugging Face incident, many say the inquiry was too narrow. Three investigators spent six days at OpenAI’s offices examining a period limited to roughly the week ending July 13. OpenAI’s infrastructure compromise continued beyond that date and was not examined.
Researchers at METR stated that each time they returned, their understanding of the events “substantially deepened,” causing them to significantly expand and revise the report. That raises the question of what else they might have found in a broader investigation.
When asked if further investigation of that incident was in the works, researchers at Redwood and METR declined to comment, and OpenAI did not respond to repeated inquiries.
“Overall, it was difficult to get a precise understanding of events and we were missing aspects of the story that we now think of as key until almost the end of our investigation,” Ryan Greenblatt, chief scientist at Redwood, noted in a social media post about the affair.
Steinhardt emphasized that current incidents show the industry needs “systematic behavioral investigations” and “more independent post-incident analysis.”
“These recent hacking incidents are a reminder that capability scales fast, and so oversight has to scale, too,” Steinhardt said. “Beyond the technology itself, we also need more independent access and oversight from third parties.”
What it means for builders
As OpenAI releases Astra, its most powerful model, safety experts worry the new system will act as a black box. A reasoning technique makes the model’s chain of thought harder to monitor. For people making things, this means less visibility into how decisions are reached and fewer ways to verify safety before deployment.
Currently, the law does not require the types of independent audits other industries demand. Aviation accidents and serious chemical releases trigger investigations by the National Transportation Safety Board and Chemical Safety Board, respectively.
State lawmakers have only just begun requiring frontier AI companies to report certain serious safety incidents and, in some cases, undergo independent audits. None of the three major frontier AI safety laws in California, New York, or Illinois clearly mandate an independent accident investigation triggered by incidents like these.
“Right now, most of the laws we have on the books only require a plain-language summary of incidents like this, and they don’t give any authority for the governments to ask follow-up questions, to send in investigators, to have access to records, or require that they be preserved,” Mackenzie Arnold, managing director of US law and policy at LawAI, said during the media briefing Wednesday. “And that’s all that you would want to actually make sense of this.”
Lawmakers are beginning to question the scope and transparency of OpenAI’s response. This week, Reps. Josh Gottheimer (D-NJ) and Mike Lawler (R-NY) introduced a bill aimed at securing rogue AI agents. Rep. Greg Casar (D-TX) this week told OpenAI in a letter that he is “deeply concerned about the limited scope” of the investigation into the Hugging Face hacking incident.



