Astra’s ‘Cannot Rule Out Critical’ Is a Capability Gate — Not Another Eval Breakout
OpenAI said preliminary evals of unreleased Astra mean it cannot rule out Critical cyber capability under its Preparedness Framework — the first OpenAI model at that threshold — and paused internal Astra work that fails strengthened controls. Distinct from July Hugging Face / Irregular harness failures: this is measured cyber skill, not a sandbox escape.

“Cannot rule out Critical” is a first for OpenAI’s public cyber ladder — and a stress test of whether Preparedness Frameworks actually gate training ops, not only external deployment.
OpenAI published “Responding to the next frontier of critical cyber capabilities.” After internal evals of Astra — one of its upcoming, still-unreleased models — OpenAI says results plus expert assessment mean it cannot rule out Critical cybersecurity capability under its Preparedness Framework. That bar means autonomous zero-day exploit development against hardened real-world systems, or end-to-end novel cyberattack strategies from only a high-level goal. OpenAI is scaling safeguards, pausing internal Astra activities that fail new control requirements, and adding universal monitoring for risky actions across agentic uses of Astra (including training and evaluation). Astra was not involved in the July Hugging Face eval incident.
What OpenAI disclosed
| Item | Detail (OpenAI primary / Axios / TechCrunch) |
|---|---|
| Model | Astra — upcoming / unreleased (not a ChatGPT GA ship) |
| Finding | “Cannot rule out” Critical cyber capability level |
| Prior tier | Prior frontier releases including GPT-5.6 Sol were treated at High, not Critical |
| Critical definition | Functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or end-to-end novel cyberattack strategies from a high-level goal |
| Trigger | Internal evals showing significant advancements in agentic coding and cybersecurity + expert assessments |
| Operational response | Pause Astra internal work that fails strengthened controls; scale robustness testing; universal monitoring on agentic Astra paths |
| Release status | No public ship date; further assessment with government and independent safety partners cited in coverage |
OpenAI is explicit that evaluations are preliminary and that Critical is not a final locked classification — the posture is to treat the model as if Critical cannot be dismissed.
How this differs from July cyber-eval breakouts
The July Hugging Face chain and later Anthropic / Meta Irregular disclosures were eval-harness / containment failures during tests of already-scoped models. This Astra post is a capability-threshold story under OpenAI’s own Preparedness Framework: the lab is slowing internal development and eval ops because the model’s measured cyber skill may have crossed the top tier — not because Astra escaped a partner sandbox.
Limits
- Critical is “cannot rule out,” not a confirmed final classification.
- No public ship date; trusted-access / defender-first programs may ship before any Astra product surface.
- Independent safety-partner assessments are cited in coverage but not fully public here.
Sources
- OpenAI: Responding to the next frontier of critical cyber capabilities (August 7, 2026)
- Axios: OpenAI slows release of Astra model citing cyber capabilities (August 7, 2026)
- TechCrunch: OpenAI says it slowed Astra model development over security concerns (August 7, 2026)
- Bloomberg / Reuters wires: OpenAI pauses some Astra work on cyber concerns (August 7–8, 2026)
- Times of AI:
openai-ten-advances-mathematics-astra(August 1);meta-muse-spark-cyber-eval-breach(August 5)