Create
OpenAI says Astra reaches critical cybersecurity level and can exploit zero-days

OpenAI says Astra reaches critical cybersecurity level and can exploit zero-days

Arkadiy Andrienko

OpenAI says its Astra model has become the first system the company has classified at its Critical cybersecurity capability threshold under the Preparedness Framework. The unreleased model can reportedly identify previously unknown security flaws and turn them into working exploits without step-by-step human guidance.

The company also says that, with the right tools and access, Astra can operate across heavily protected systems and carry out attacks without a person directing every step. In internal testing, Astra achieved a perfect 100% score on ExploitBench, which measures the ability to create exploits from known vulnerabilities. OpenAI says the model also discovered and used two zero-day vulnerabilities as part of an exploit chain.

Comparison of Astra and GPT-5.6 Sol models
Comparison of Astra and GPT-5.6 Sol models

The company says Astra went further in another test as well: it built a browser-compromise chain that escaped the sandbox and executed commands on the host. OpenAI also says the model found multiple vulnerabilities in a hardened operating system and combined them into a local privilege-escalation chain from an unprivileged user to root.

Because of these capabilities, OpenAI delayed parts of Astra's development while it strengthened protections against cyber misuse and unauthorized model actions. The company says Astra now refuses 91.5% of disallowed cyber requests, compared with 59% for GPT-5.6 Sol.

OpenAI plans to make Astra available soon, but access to its most advanced cybersecurity capabilities will be limited at first to a small group of alpha testers and then to Daybreak Blue. "

It is the first model we are designating at this level.
— OpenAI, developer

"

If a model can independently find and exploit zero-days, how much control is enough before it is released?

    About the author
    Comments0