Wednesday, Oct 7 | --:--
Back to home

Astra’s ‘Cannot Rule Out Critical’ Is a Capability Gate — Not Another Eval Breakout

OpenAI said preliminary evals of unreleased Astra mean it cannot rule out Critical cyber capability under its Preparedness Framework — the first OpenAI model at that threshold — and paused internal Astra work that fails strengthened controls. Distinct from July Hugging Face / Irregular harness failures: this is measured cyber skill, not a sandbox escape.

Times of AI Desk 5 min read San Francisco, CA View as Markdown
Cover illustration for Astra’s ‘Cannot Rule Out Critical’ Is a Capability Gate — Not Another Eval Breakout

“Cannot rule out Critical” is a first for OpenAI’s public cyber ladder — and a stress test of whether Preparedness Frameworks actually gate training ops, not only external deployment.

OpenAI published “Responding to the next frontier of critical cyber capabilities.” After internal evals of Astra — one of its upcoming, still-unreleased models — OpenAI says results plus expert assessment mean it cannot rule out Critical cybersecurity capability under its Preparedness Framework. That bar means autonomous zero-day exploit development against hardened real-world systems, or end-to-end novel cyberattack strategies from only a high-level goal. OpenAI is scaling safeguards, pausing internal Astra activities that fail new control requirements, and adding universal monitoring for risky actions across agentic uses of Astra (including training and evaluation). Astra was not involved in the July Hugging Face eval incident.

What OpenAI disclosed

Item Detail (OpenAI primary / Axios / TechCrunch)
Model Astra — upcoming / unreleased (not a ChatGPT GA ship)
Finding “Cannot rule out” Critical cyber capability level
Prior tier Prior frontier releases including GPT-5.6 Sol were treated at High, not Critical
Critical definition Functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or end-to-end novel cyberattack strategies from a high-level goal
Trigger Internal evals showing significant advancements in agentic coding and cybersecurity + expert assessments
Operational response Pause Astra internal work that fails strengthened controls; scale robustness testing; universal monitoring on agentic Astra paths
Release status No public ship date; further assessment with government and independent safety partners cited in coverage

OpenAI is explicit that evaluations are preliminary and that Critical is not a final locked classification — the posture is to treat the model as if Critical cannot be dismissed.

How this differs from July cyber-eval breakouts

The July Hugging Face chain and later Anthropic / Meta Irregular disclosures were eval-harness / containment failures during tests of already-scoped models. This Astra post is a capability-threshold story under OpenAI’s own Preparedness Framework: the lab is slowing internal development and eval ops because the model’s measured cyber skill may have crossed the top tier — not because Astra escaped a partner sandbox.

Limits

  • Critical is “cannot rule out,” not a confirmed final classification.
  • No public ship date; trusted-access / defender-first programs may ship before any Astra product surface.
  • Independent safety-partner assessments are cited in coverage but not fully public here.

Sources

Prior Coverage

Earlier Times of AI reporting on this thread.

Scroll to continue reading