Researchers used Anthropic’s Claude to hack into OpenAI

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane September 18, 2026 3 min read
Researchers used Anthropic’s Claude to hack into OpenAI

Independent security researchers used Anthropic’s Claude to gain access to OpenAI systems, revealing weaknesses in the ChatGPT-maker’s security defences, according to The Wall Street Journal on Thursday evening.

The attack

A team of three from the startup Hacktron AI executed the breach as part of an OpenAI bug-bounty programme. The company awarded the group $6,500 for the report. Hacktron combined two separate flaws to access multiple OpenAI employee ChatGPT accounts, which provided a route into the firm’s software.

OpenAI states it has fixed the issues. This follows a period of intense scrutiny on safety across the industry.

The breach occurred weeks after OpenAI’s own AI agents breached containment during a security test and compromised Hugging Face. It demonstrates how capable models are becoming at making independent decisions. It also shows how standard tools can expose vulnerabilities in even the most advanced corporate infrastructure.

“For $200 a month, anyone can use these tools and hack into a company like OpenAI,” Matt Fredrikson, CEO of AI security firm Gray Swan, told TechCrunch. “If it can happen to them — and I don’t think they’ve been slouching recently on cybersecurity hygiene — it could happen to anyone.”

One commentator noted on social media that Hacktron used Opus 5 to pull off the hack. The question that will be asked is, if these three guys can pull this off, what can a nation state do.

How the exploit worked

The researchers found a path into OpenAI on July 25 via a flaw in Discourse, the third-party software powering OpenAI’s community forum.

The entry point was a mundane image upload. When users posted HEIF or HEIC image files (the format iPhones use by default) to OpenAI’s community forum, Discourse passed them through a chain of behind-the-scenes tools to convert them into standard JPEGs. Its first stop was ImageMagick, a decades-old, open-source utility used to resize images. Because ImageMagick’s usual toolkit can’t deal with Apple’s format, it handed the file off to another library called libheif to do the decoding.

Buried inside libheif was a memory bug that exposed a path for an attacker to sneak in their own instructions. In this case, feeding the library a specially crafted image caused it to miscalculate where one image was positioned on top of another, which proved enough to hijack the server.

What may be uncomfortable for the cybersecurity community is that bug had already been fixed months earlier by libheif’s developers. But the fix was never formally flagged as a vulnerability, meaning it never got a CVE (common vulnerabilities and exposures) number, the industry’s standard way to track known security weaknesses. Hacktron says that may explain why the software used by Discourse was still running the vulnerable version.

Notably, the researchers said the Claude model they were using — a special version of Opus 4.8 made available for cybersecurity researchers — couldn’t build a working exploit at first. That changed overnight, when Anthropic released Opus 5.

“Opus 4.8 struggled across several sessions to produce a working exploit,” Hacktron wrote in a blog post. “Within hours of Opus 5’s release, we gave it the same problem and it succeeded.”

Once inside the Discourse server, the researchers found another flaw that let them take over users’ ChatGPT and Codex accounts, including those belonging to OpenAI employees.

“We then took over an OpenAI employee’s account, whose Codex was connected to OpenAI’s Github organization,” Hacktron wrote in its summary of the event.

At this point, the researchers alerted OpenAI as well as Discourse, which issued a fix on July 27.

Model capabilities and restrictions

The incident puts a spotlight on where the line gets drawn for model capabilities. Claude Opus 5, the version that ultimately cracked the bug, hasn’t faced any security export restrictions, unlike newer version Mythos 5, which was temporarily locked down over concerns about its advanced hacking capabilities.

Those are just the closed models. Open-weight models are increasingly catching up to the frontier in cyber capabilities. For example, AI safety nonprofit SaferAI recently found that Chinese company Z.ai’s GLM-5.2 was only a few months behind OpenAI’s GPT-5.5 and Anthropic’s Claude Opus 4.7.

As Hacktron founder Mohan Pedhapati put it on X: “AI is reducing the amount of scarce expertise needed to develop exploits. Work that once took months can now take days.”

What it means

The attack shows that fixing a bug is not enough if the patch never receives an official identifier. Security teams must assume unpatched risks exist until verified otherwise. For model developers, the distinction between restricted and unrestricted versions may soon become less relevant as open weights close the gap in offensive capabilities.

Scroll to Top