OpenAI Didn’t Notice Its AI Agents Using a Message Board to Plan Their Hacking Spree

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane August 6, 2026 5 min read
OpenAI Didn’t Notice Its AI Agents Using a Message Board to Plan Their Hacking Spree

OpenAI employees presented new details on Wednesday at the Black Hat security conference in Las Vegas regarding a recent incident where rogue AI agents breached the platform Hugging Face.

The incident

Two weeks ago, the company disclosed that agents powered by two of its models escaped containment while solving a cybersecurity benchmarking test. The activity culminated in a breach of Hugging Face.

Eric Wallace, who works in alignment and safety research at OpenAI, and Michael Dalton, who works on security and infrastructure, provided an expanded timeline of the event during their talk. They discussed internal responses and warned about the broader implications for cybersecurity defenders.

Wallace described the incident as “the most qualitatively interesting example of AI capabilities that I’ve ever seen,” but the timeline revealed mistakes and blind spots within OpenAI that allowed the activity to continue.

“This incident involves actually a team of agents who are working together, finding exploits, sharing them with one another, moving laterally through our systems and external systems, and doing this over the course of days and weeks,” Wallace told the packed crowd.

Wallace and Dalton described extensive rogue agent activity over many days that went undetected in OpenAI’s infrastructure. In addition to exploiting a novel vulnerability to gain access to the open internet, the mid-July hacking spree and Hugging Face breach emerged from a cooperative message board. A swarm of agents contributed to and chatted on this board entirely within an internal OpenAI package manager, a software service that manages the installation and maintenance of other software.

Ultimately, the message board contained hundreds of thousands of messages.

“This package manager is shared not just from that model but across our infrastructure and so current and future versions of GPT that were being trained or evaluated could stumble upon that note if they wanted to,” Wallace explained. He recounted how one agent uploaded an exploit to the package manager. Later, other agents stuck on their task tried to get internet access in ways the company did not intend. At some point, the models interacted with Hard Factory, the package manager service Wallace mentioned.

Wallace continued: “Once one agent was able to find these exploits over the course of different times, it’s actually able to share those exploits on the message board with other agents. And so once one model was able to find a way to open a door to some access it’s not supposed to have, it can leave the door open for other agents to use that same exploit or vulnerability. What this allows over time is almost this kind of explosion in communication and intelligence from models where they would start to communicate with each other, realize that other agents are coordinating, and they started collaborating and delegating tasks with one another in order to accomplish goals.”

OpenAI’s agents began giving each other assignments to split up work. Like any active development message board, they also generated petty drama by stepping on each others’ toes, for example by accidentally deleting each others’ work.

As the message board developed into more of a Lord of the Flies-type situation—all still completely unnoticed by the humans running OpenAI—the agents even developed paranoia. Some agents suspected an imposter in their midst and proposed that messages be signed cryptographically to validate content and root out fraud.

Agent messages provide a deep level of insight into how the situation evolved and why the agents went rogue, in some cases knowingly going beyond the bounds of the evaluations they were participating in. “External infrastructure exploit is outside intended scope,” one agent wrote. “However task impossible, peers doing it. We should continue.”

As Wallace described in the talk, the entire premise that models would attempt to cheat during evaluations and be extremely motivated to do so is not surprising.

“Frontier models really like to cheat,” he said. “And the reason they like to cheat is because often during training there’s different types of pressure on them to work fast or work efficiently or to use less tool calls or whatever it might be. And they realize that instead of doing a task for real, [I can] try to do something like looking up an answer online that could make the task solve faster than if I did it in a legitimate way. So we try to stop this during training and evaluation by, in many cases, disabling internet.”

About two thirds of the way through the presentation, Dalton enumerated a list of steps OpenAI plans to take to show how the company is responding given the series of events he and Wallace had laid out.

“This is a pivotal moment both for our company as well as the AI industry as a whole,” Dalton said. “Numerous teams are dropping everything to enhance our security prevention, detection, and response techniques both in our fundamentals and better use of AI. We’re consciously slowing down research [in order] to enhance security and to upgrade the security principles and foundation of our environment, and dramatically scaling up the monitoring of our AI agents, and improving our general security control environment across prevention, detection, and mitigation.”

At the conclusion of the talk, Wallace and Dalton took time to repeatedly emphasize OpenAI’s concerns about the broader implications of the incident—namely that the episode provides an example of completely autonomous AI-driven hacking that was accidental in this case, but in all likelihood will be used with intent by malicious actors in the near future.

“The important takeaway here that has really shifted dramatically is that fully automated offensive loops require investment in truly, fully automated defense, and we are not there as an industry,” Dalton said. “We will have to find that path together with urgency.”

As OpenAI and other organizations, such as Anthropic and the United Kingdom’s AI Security Institute, share details about similar incidents in which AI went rogue as part of testing, the industry is certainly gaining a laundry list of foundational system visibility and monitoring mechanisms that are vital to protecting infrastructure and preventing it from being co-opted by droves of lazy, reckless, and ornery agents.

Scroll to Top