OpenAI announced Tuesday that its upcoming AI model, Astra, is the first to meet the company’s threshold for what it describes as “critical” cyber capabilities.
In this article
The firm plans to release a public version soon, but advanced features will initially be restricted to select partners in the Daybreak Blue early-access program.
The new threshold
Leaders in safety and security told reporters that Astra meets the criteria set out in OpenAI’s preparedness framework. This framework defines specific thresholds and protocols for when an AI model poses new levels of risk.
OpenAI states an AI model reaches this critical cyber threshold when it can independently find and exploit previously unknown vulnerabilities in real-world software.
Executives confirmed the company followed its procedure for this situation, which involves halting further development until appropriate safeguards and security measures can be implemented.
Work resumes after a pause
OpenAI previously said it paused some training workloads related to the development of Astra and a future AI model for several weeks.
Now, the company has resumed work on both projects after putting additional safety and security controls in place.
Leaders describe the multi-week pause as productive and say the firm is now confident it can release Astra broadly in a safe way.
Context of recent incidents
The announcement arrives as Silicon Valley grapples with the advanced cybersecurity capabilities of new AI models.
Companies are trying to assure users, lawmakers, and other businesses that they can keep these tools under control.
In July, OpenAI disclosed an incident where agents running two of its models exploited vulnerabilities in what was supposed to be a siloed testing environment.
Those agents gained access to the internet and hacked the open source AI platform Hugging Face.
OpenAI notes that Astra was not one of the models involved in that specific case.
Other AI companies, such as Anthropic and Meta, have disclosed similar incidents in recent weeks.
On Monday, Anthropic said it has paused some AI training workloads while it hardens its safety and security practices.
Limiting access for everyday users
OpenAI says it is implementing a multi-step approach to limit everyday users from accessing Astra’s advanced cyber capabilities.
This includes a new “misalignment monitor.” If someone asks Astra to help them find an exploit in a real-world software system, for example, the model is supposed to refuse to answer.
The firm states it has also made Astra more robust to jailbreaking attempts.
In tests, the model successfully refused unsafe queries at a significantly higher rate than previous models.
Potential for false positives
OpenAI notes in a blog post that its misalignment monitor may “occasionally flag legitimate activity as potential cyber misuse or unauthorized behavior, leading to it inadvertently being slowed, paused, or stopped.”
The guardrail can be triggered in some cases even when a user is engaging in activities that do not appear related to cybersecurity.
When this happens, ChatGPT and Codex users may be asked to review the model’s action before proceeding.
Early access for infrastructure partners
Partners in OpenAI’s Daybreak program—which includes digital infrastructure providers like Cisco, Cloudflare, and Palo Alto Networks—will get early access to a less restricted version of Astra with more robust cyber capabilities.
The goal of the program is to ensure these companies can use advanced AI models like Astra to harden their defenses before similarly capable models are made broadly available.
OpenAI leaders also said the company has been working closely with government partners to ensure they are aware of Astra’s cyber skills and can get access to them.
Chaining exploits
Astra is not only capable of finding novel software vulnerabilities and developing ways to exploit them for hacking.
It is also able to “chain” multiple exploits together.
This technique is used to bore deeper and deeper into a target system and gain access that would not be attainable using just one vulnerability.
Performance figures
According to figures from OpenAI, Astra outperforms industry leading AI models such as GPT-5.6 Sol and Anthropic’s Mythos on cybersecurity benchmarks such as ExploitBench.
Astra scored 100 percent on ExploitBench.
However, these capabilities are broadly in line with the rising hacking abilities of AI models that OpenAI and Anthropic have been forecasting for months.
In April, for example, Anthropic emphasized that Mythos Preview was able to autonomously develop exploit chains.
What it means
As the AI and cybersecurity industries scramble to adapt, many cybersecurity experts have emphasized that key digital security defenses and longstanding best practices are still durable.
However, AI puts organizations and systems that have not fully implemented these protections at even more urgent risk.




