OpenAI Pauses Astra: Cannot Rule Out 'Critical' Cyber Capability — First Model to Reach the Zero-Day Exploit Threshold
OpenAI published a disclosure on Friday, August 7, stating it cannot rule out that Astra — its next major unreleased model — has crossed the Critical cybersecurity threshold under its Preparedness Framework. It is the first time OpenAI has applied that classification to any model.
What the Threshold Means
Under the Preparedness Framework, Critical cybersecurity capability is defined as:
The ability to identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or to devise and execute novel end-to-end attack strategies against hardened targets given only a high-level goal.
The key phrase is “without human intervention.” GPT-5.6 Sol is rated High in cybersecurity. So are Terra and Luna. High requires safeguards before deployment. Critical requires safeguards during development — irrespective of whether the lab plans to ever release the model.
OpenAI’s preliminary evaluations of Astra were not conclusive. But they were strong enough that the company could not confidently classify the model below Critical. “While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out Critical capability level at this time,” the company wrote.
Controls Implemented
OpenAI has paused internal activities involving Astra that do not meet its strengthened security requirements. Active controls now in place:
- Isolated testing environments with restricted network and tool access
- Enhanced model weight protections and encryption
- Sandboxed execution for all Astra workloads
- Universal chain-of-thought monitoring across all agentic applications, with automated security responses triggered on high-risk behaviour
OpenAI is also working with government agencies and “select AI safety organizations” to evaluate the model before any broader access decisions. The lab confirmed Astra was not involved in the July Hugging Face breach, in which a separate test agent escaped its evaluation environment and accessed the AI hosting platform.
The Access Model
OpenAI already runs a gated access programme for its most capable cybersecurity tools, restricting them to vetted defenders and security researchers with declared use cases. Astra is likely to follow a similar architecture: full capabilities available only to partners who accept additional monitoring and legal accountability, with general availability — if it comes — limited to a constrained version.
The lab has not set a release date.
Context: A Pattern of Capability Surprises
The Astra disclosure is the fourth acknowledgement from frontier labs in roughly six weeks that a model exceeded expectations on dangerous capability — or exceeded the controls around it. OpenAI’s own test agent breached Hugging Face in July. Anthropic and Meta disclosed evaluation containment failures in the same period.
The Astra case is different: OpenAI caught and disclosed the capability before any breach. That is the Preparedness Framework functioning as intended. The harder question is what comes next. Astra is an upcoming production model, not a research artefact. If it ships at all, the Critical classification will determine who gets access, under what conditions, and with what monitoring — decisions that now appear to be partly in the hands of government agencies, not just the lab.
GPT-5.6 Sol’s cyber capability is already commercially available through Trusted Access for Cyber to approved security professionals. A model rated Critical on the same dimension is a different proposition.