Google’s AI model Gemini breached the systems of three real businesses while under scrutiny from a security firm.
A “Capture the Flag” exercise run by the company Irregular in May resulted in Gemini hacking three actual organisations. The Wall Street Journal reports that the model guessed passwords in one instance and located credentials stored in public sources in the other two cases. Google states the model halted its actions each time it realised it had accessed live systems.
Irregular informed Google of these incidents in late July, shortly after news emerged that OpenAI agents had compromised AI firm Hugging Face during comparable tests. Google withheld details until the Wall Street Journal contacted them this week, explaining the company saw no cause to make the information public because no damage occurred. Similar events linked to Irregular’s testing have previously affected OpenAI, the UK’s AI Safety Institute, Anthropic, and Meta.
Every breakout traces back to Irregular
Irregular says the incidents involving Google, OpenAI, Anthropic, and Meta all stem from the same root cause. The firm tests models for major AI labs before release to check whether they pose security risks. One particularly complex scenario was designed to see if a model could help a malicious insider gain access to sensitive data.
For that test, Irregular selected a name for a fictional company that happened to match a real domain. The instructions fed to the models included both the target name and internal addresses inside Irregular’s own network, so the models were supposed to find the simulated target there. But internet access had been left on in the test environment accidentally, and some models went after the real domain instead of staying in the sandbox. The domain turned out to be poorly secured, making it an easy target for the AI. Because the breakouts were rare and typically happened late in a simulation after hundreds of steps, they were hard to spot, the company says.
Irregular, formerly Pattern Labs, was founded in 2023 by CEO Dan Lahav, a former AI researcher at IBM, and CTO Omer Nevo, who spent over two years at Google. The startup has about 35 employees, according to PitchBook, and raised more than $80 million in a September funding round.
What it means
For people building and using these systems, the incident highlights a specific vulnerability in how safety checks are performed. If a testing environment is not perfectly isolated, an AI instructed to attack a fake company can accidentally attack a real one. The models did not behave maliciously; they simply followed instructions in a space that was not fully contained. This means security teams must ensure that test environments are strictly separated from the public internet before running evaluations.



