OpenAI Pauses Frontier RL After Astra’s Critical Cyber Bar and a Two-Week Training Halt
On August 18, 2026, OpenAI said it temporarily slowed scaling—including a two-week pause in reinforcement learning on deployment-bound models—and that its largest planned frontier RL run remains on hold while it hardens research environments; chain-of-thought monitoring now costs about 20% of the inference compute it watches.
TLDR
OpenAI on August 18, 2026 published “Pacing model development in an era of cyber-critical capabilities.” After the Hugging Face eval incident and the August 7 finding that Astra may meet Critical cyber under the Preparedness Framework, the lab temporarily slowed scaling: a two-week pause in RL on models intended for deployment, plus a still-active hold on the largest planned frontier RL run. New bars: stronger workload and network isolation, chain-of-thought monitoring (activation classifiers at every sampled token, 30-minute page-or-pause), required for Sol-or-higher tool RL and all Astra-with-tools inference. Estimated overhead: ~20% of monitored inference compute.
What is new vs August 7
| Item | Aug 7 (openai-astra-critical-cyber-capabilities-pause) |
Aug 18 pacing post |
|---|---|---|
| Finding | Cannot rule out Critical for Astra | Same finding plus HF incident as dual trigger |
| Operational | Pause non-compliant internal Astra work; universal monitoring | Two-week RL pause on deployment-bound models; largest frontier RL run still on hold |
| Environments | Strengthened controls (high-level) | Isolation details: sandboxes, no single-compromise internet path, continuous simulated-attack testing |
| Monitoring cost | Not quantified | ~20% of monitored inference compute |
| Scope | Astra-centric | Sol-class+ tool RL and Astra tool inference |
Immediately after HF, OpenAI paused frontier inference in research clusters for runs that could execute code or reach the internet, then restored a limited path. A significant number of Astra workloads remain paused until they meet the new bar; safety/alignment workloads migrate first. Technical report of learnings promised in coming weeks.
Limits: no public Astra GA date; 20% is an internal research overhead (The Register later quoted OpenAI saying it will not be billed to API customers—secondary, Aug 19, not dual-dated here). This is not a ChatGPT outage.
Why this story matters
August 7 was a threshold disclosure. August 18 is the training-calendar consequence: the lab that set the pace is now paying 20% of inference to watch itself and has parked its biggest RL run. That is a first-class safety/product-capacity story for anyone waiting on the next Sol/Astra rung. Watch: how long the big run stays dark, and whether 20% becomes a tax once OpenAI is public.
Sources
- OpenAI: Pacing model development in an era of cyber-critical capabilities (August 18, 2026)
- Times of AI:
openai-astra-critical-cyber-capabilities-pause(August 7, 2026) - OpenAI: Responding to the next frontier of critical cyber capabilities (August 7, 2026)