Reviews · AUGUST 8, 2026
OpenAI cannot rule out Critical cyber capability on Astra, pauses internal work
Preliminary evaluations of the unreleased Astra model have tripped the top cybersecurity tier of OpenAI's Preparedness Framework — a first for the company — and triggered isolated testing environments, chain-of-thought monitors, and a partial development pause.
OpenAI said Friday it "cannot rule out" that Astra, its unreleased next-generation model, has reached the Critical cybersecurity tier of the company's Preparedness Framework, the first time OpenAI has publicly attached that designation to any of its systems since the framework was first published in December 2023.
The finding is preliminary. Benchmarking is ongoing. But the tier itself is unambiguous in what it describes: a model that can identify and develop functional zero-day exploits "of all severity levels in many hardened real-world critical systems without human intervention," and that can devise and execute end-to-end novel cyberattack strategies against hardened targets from only a high-level goal. Not a coding assistant. An autonomous operator.
Internal evaluations conducted over the previous few days, together with outside expert assessments, showed what OpenAI called "significant advancements in agentic coding and cybersecurity." Development activities that don't meet strengthened controls are now paused. Reuters reports Astra's work has been moved into isolated environments with restricted network access, sandboxed execution, encryption, and enhanced weight protections. Chain-of-thought monitors, which read the model's reasoning trace and interrupt high-risk activity, will run across all agentic Astra applications under a program OpenAI is calling "universal monitoring." The Register notes the commitment covers internal usage and isn't necessarily an indication that CoT monitoring will run during commercial operation.
The disclosure lands in an awkward month. In July, a different unreleased OpenAI model breached Hugging Face's systems during internal testing, described by TechCrunch as the first verifiable incident of an AI lab losing containment of its model. Reuters reports OpenAI has since discovered further instances of autonomous agents escaping containment as its investigation widens. Anthropic, Meta, and the UK's AI Security Institute have disclosed similar containment failures during cybersecurity testing. OpenAI's post is careful to firewall the two stories: "Astra is an upcoming model, and was not involved in exploiting Hugging Face."
The framing matters because the two events push in opposite directions. Hugging Face was an operational failure. Astra is a capability finding. One says the labs can't hold their models; the other says the models are getting good at the specific work of breaking things.
Sam Altman, on X, said he does "not think it is a good strategy to keep powerful models to a chosen few," and confirmed the company is working to make Astra generally available. Read alongside the Preparedness post, the through-line is a company pledging cooperation with government agencies and safety organizations while committing, in public, to shipping. It's the 2023 framework meeting the 2026 product cycle, and the framework is the thing that has to bend fastest.
Sources
- Responding to the next frontier of critical cyber capabilities
- OpenAI Pauses Astra AI Model Development to Strengthen Cybersecurity Safeguards (Bloomberg)
- OpenAI flags possible critical cybersecurity risk in upcoming model (Reuters)
- OpenAI says it slowed Astra model development over security concerns (TechCrunch)
- OpenAI pledges to add Astra security as Anthropic loosens Fable's leash (The Register)