# OpenAI Parks Its Biggest Frontier RL Run — and Pays 20% to Watch Itself

Times of AI Desk · 2026-08-18 · Research

[https://timesof.ai/2026/08/openai-pacing-cyber-critical-rl-pause](https://timesof.ai/2026/08/openai-pacing-cyber-critical-rl-pause)

> After Astra’s Critical cyber finding and the Hugging Face eval incident, OpenAI temporarily slowed scaling — including a two-week RL pause on deployment-bound models — and left its largest planned frontier RL run on hold. Chain-of-thought monitoring now costs about 20% of the inference compute it watches.

August 7 was a **threshold disclosure**. August 18 is the **training-calendar** consequence: the lab that set the pace is now **paying 20% of inference** to watch itself and has **parked its biggest RL run**.

**OpenAI** published **“Pacing model development in an era of cyber-critical capabilities.”** After the **Hugging Face** eval incident and the **August 7** finding that **Astra** **may meet Critical** cyber under the Preparedness Framework, the lab **temporarily slowed scaling**: a **two-week pause in RL** on models **intended for deployment**, plus a still-active hold on the **largest planned frontier RL run**. New bars: stronger **workload and network isolation**, **chain-of-thought monitoring** (activation classifiers at **every sampled token**, **30-minute** page-or-pause), required for **Sol-or-higher** tool RL **and** all **Astra-with-tools inference**. Estimated overhead: **~20%** of monitored inference compute.

## What is new vs August 7

| Item | Aug 7 (Critical cyber pause) | Aug 18 pacing post |
|------|------------------------------|---------------------|
| **Finding** | Cannot rule out **Critical** for **Astra** | Same finding **plus** HF incident as dual trigger |
| **Operational** | Pause **non-compliant internal Astra** work; universal monitoring | **Two-week RL pause** on **deployment-bound** models; **largest frontier RL run still on hold** |
| **Environments** | Strengthened controls (high-level) | Isolation details: sandboxes, no single-compromise internet path, continuous simulated-attack testing |
| **Monitoring cost** | Not quantified | **~20%** of monitored inference compute |
| **Scope** | Astra-centric | **Sol-class+** tool RL **and** Astra tool inference |

Immediately after HF, OpenAI **paused frontier inference** in research clusters for runs that could **execute code or reach the internet**, then restored a **limited** path. **A significant number of Astra workloads remain paused** until they meet the new bar; **safety/alignment workloads migrate first**. Technical report of learnings promised in **coming weeks**.

## Limits

- No public Astra GA date.
- 20% is an **internal research** overhead (The Register later quoted OpenAI saying it **will not** be billed to API customers — secondary, Aug 19).
- This is **not** a ChatGPT outage.

## Sources

- [OpenAI: Pacing model development in an era of cyber-critical capabilities (August 18, 2026)](https://openai.com/index/pacing-model-development-cyber-capabilities/)
- Times of AI: `openai-astra-critical-cyber-capabilities-pause` (August 7, 2026)
- [OpenAI: Responding to the next frontier of critical cyber capabilities (August 7, 2026)](https://openai.com/index/responding-next-frontier-critical-cyber-capabilities/)
