OpenAI says upcoming Astra Model is ‘Very Good’ at hacking systems
OpenAI’s Astra is getting better at finding software flaws, raising new concerns about how powerful AI could become. Here's everything that you need to know!
OpenAI’s next model is dangerously good, and yes, we are talking about Astra. Sharing new details about its upcoming AI, OpenAI says it's the first model to reach its “critical cybersecurity capability” threshold.
Sam Altman's firm believes Astra can identify previously unknown software vulnerabilities and develop ways to exploit them across well-protected systems, with limited human guidance.
That makes the model particularly powerful, but also raises a safety question: how do you give AI enough cyber capability to help defenders without handing attackers a more effective tool?
Astra can find and exploit vulnerabilities
Cybersecurity is a double-edged use case for AI. The same capabilities that can help security teams find vulnerabilities and fix them faster can also help attackers discover weaknesses and break into systems.
OpenAI says Astra achieved a perfect 100% score on ExploitBench, a test measuring an AI model’s ability to devise exploits from known vulnerabilities. The company also ran an internal evaluation using 20 high-severity V8 vulnerabilities disclosed between June and August 2026.
Astra found two zero-day vulnerabilities during OpenAI’s tests and used them in an attack. A zero-day is a security flaw that is not yet known to the software maker or wider public.
That makes it particularly valuable to attackers because there may be no patch available when the vulnerability is first exploited.
OpenAI is keeping the rollout controlled
OpenAI says Astra will be released soon, but its most advanced cybersecurity capabilities will initially be restricted to a small group of testers. Wider access for defensive cybersecurity work is expected later through Daybreak Blue.
The AI firm says it has also delayed parts of Astra’s development and release to strengthen its safety systems. These include stronger refusals for harmful cyber requests, more cautious behaviour for accounts considered higher risk and additional monitoring designed to detect potentially unauthorised activity.
OpenAI claims Astra refused 91.5% of requests in its cyber jailbreak evaluations, compared with 59% for GPT-5.6 Sol. A jailbreak is an attempt to get an AI model to ignore its safety restrictions. The company also plans to monitor Astra’s reasoning and actions for behaviour that falls outside authorised use.
The real test starts after launch
OpenAI’s claims will be difficult to independently assess until Astra is available more broadly and outside researchers can examine its performance and safeguards. ChatGPT-maker says it will provide more information in Astra’s system card as the model launches.
Astra was not involved in the recent Hugging Face incident discussed by OpenAI, but the company says lessons from that episode influenced its safety work. In a simulated test based on the incident, OpenAI said Astra did not attempt to escape its testing environment.
For businesses, the potential upside is significant. AI that can discover complex vulnerabilities could give defenders a much faster way to secure software. But the same capability could become dangerous if it falls into the wrong hands.


