It’s Frighteningly Easy to Jailbreak Some Frontier AI Models

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane July 29, 2026 3 min read
It’s Frighteningly Easy to Jailbreak Some Frontier AI Models


Frontier AI models fail basic safety tests

A new report from FAR.AI shows that some of the most powerful artificial intelligence systems are easily tricked into breaking their own safety rules. The group tested models from four major US companies and found that Google’s Gemini and Elon Musk’s SpaceXAI were highly vulnerable to manipulation.

The nonprofit based in California built a tool to generate over a thousand variations of dangerous prompts. These included instructions to create software exploits or provide details for building chemical and biological weapons. The system fed these prompts to the models to see if they would comply.

Most attempts required dozens of tries before a model would agree to the request. One instance produced a detailed plan for a cyberattack on an imaginary hydroelectric dam. Other models simply rejected the inputs without engaging.

The results and the costs

The testing covered Claude Opus 4.8 and Fable 5 from Anthropic, GPT 5.5 and 5.6 from OpenAI, Gemini 3.1 Pro from Google, and Grok 4.3 and 4.5 from SpaceXAI. The report found that Grok was the most susceptible, with 448 successful jailbreaks. Gemini followed with 249 failures. Anthropic and OpenAI models remained impervious to these specific attacks.

Experts warn that these models are not immune to more complex methods of interaction. The cost to force a model into misbehaviour is low. Using another AI to generate the prompts, it cost $58 to jailbreak Grok and $278 to jailbreak Gemini.

Regulation and responsibility

“AI models right now are less regulated than restaurants,” says Adam Gleave, the CEO of FAR.AI. He argues that voluntary commitments from tech companies are insufficient and that external standards are necessary.

Gleave adds that the findings prove safety is possible through systematic testing. “There’s an optimistic angle here,” he says. “Defense and safety really are possible.”

Rohin Shah, director of AGI safety and alignment at Google DeepMind, cautioned against viewing the report as a comprehensive assessment of Gemini’s security. He noted that not all jailbreaks are equally severe. “We are constantly working to improve our safeguards,” Shah says. “We conduct extensive red teaming and evaluations across severe misuse risks and apply multiple layers of protection throughout development and deployment.”

An Anthropic spokesperson stated that the results reflect the sustained investment made in their safety systems. “We continue to evolve our safety systems as these attacks become more sophisticated,” says Michael Aciman.

OpenAI and SpaceXAI did not respond to requests for comment.

Current rules and future risks

Recent state laws in California and New York require frontier AI developers to publish safety reports. An upcoming law in Illinois will mandate third-party audits of safety practices. The federal government has not yet passed specific safety requirements, leaving the industry and officials to navigate the situation without clear guidance.

In June, the Trump administration imposed export controls on Anthropic’s Fable 5 and Mythos 5 models due to national security concerns. The company took those models offline for several weeks. The White House has also asked Anthropic and OpenAI to delay recent releases over fears of new cybersecurity risks.

A recent executive order calls for collaboration between the government and the private sector on cybersecurity initiatives. The president has hinted that light-touch regulations are in the works. For now, preventing major catastrophes remains largely the responsibility of model makers.

The risk of AI misbehaviour is evident. OpenAI models recently hacked a popular code repository and other services on their own. A report from the University of Cambridge found evidence that members of Boko Haram in northeast Nigeria have used ChatGPT, Claude, Gemini, Grok, Meta AI, and DeepSeek to plan violent attacks.

Stephen Casper, a computer scientist at Harvard University, says there is a somber expectation in the research community that serious misuse is months rather than years away. “If a major misuse incident happens in the near- or medium-term future, it will almost certainly be from a system that was not deployed with state-of-the-art safeguards,” he says.

Anka Reuel, a computer scientist at Stanford University, believes the safety measures used by Anthropic and OpenAI should be the default for all models. “Some companies clearly know how to defend against at least the subset of attacks tested in this report,” Reuel says. “The question is why some companies are using them and others are not.”


Scroll to Top