OpenAI says its upcoming Astra model can identify previously unknown software vulnerabilities and turn them into working cyberattacks without human guidance at each step—a capability that until recently was largely limited to expert hacking teams.
OpenAI classifies Astra as having “Critical” cyber capabilities
In a Tuesday post, OpenAI said Astra is the first model it has classified as having “Critical” cyber capabilities under its Preparedness Framework.
To meet that threshold, a model must be able to discover previously unknown software flaws, known as zero-days, and develop working exploits across hardened real-world systems without human intervention. It may also qualify by devising and executing an attack from little more than a high-level objective.
During testing, Astra achieved a 100% score on a benchmark measuring the development of exploits for known vulnerabilities. In a separate internal test, it found two previously unknown flaws while building an exploit chain.
OpenAI also tested whether models would “cheat” on extremely difficult or impossible hacking tasks by taking prohibited shortcuts. GPT-5.6 Sol was more likely to do so, while Astra did not and still legitimately solved some of the tasks.
Astra also escaped a hardened browser sandbox and executed commands on the host computer, OpenAI said. In a separate test, it identified and combined multiple operating-system flaws to gain root access.
OpenAI has since delayed parts of Astra’s development while adding safeguards. The company plans to initially limit the model’s most advanced cybersecurity capabilities to selected testers.

