OpenAI said that the OpenAI-Hugging Face evaluation incident and preliminary evidence that its then-unreleased Astra model may meet the Critical cybersecurity threshold led it to slow scaling, including a two-week pause in reinforcement learning training on deployment models. It describes safeguards applied during training (monitoring, alignment evidence and security isolation of research environments) and says it will evolve the Preparedness Framework accordingly.
Why it matters
It is a public case of a lab applying its Critical cyber threshold to development itself, including isolating its own training environments.
Key facts
As stated in the sources, with where to find them.
- Included a two-week pause in RL training on models intended for deployment; the largest planned frontier RL run remains on hold.Opening section
- Security measures include stronger workload sandboxes and network isolation so a single compromise cannot enable unauthorized internet access.Security measures
- Monitoring overhead is estimated at roughly 20% of the inference compute being monitored; OpenAI aims to issue an alert within 30 minutes after concerning activity is surfaced by its monitoring system.Monitoring section
Findings that cite this record
No tracked finding cites this record yet.
Key questions this bears on
- Can AI agents defend and oversee systems on their own?Not yet. Agents are weak on realistic defensive benchmarks and monitors can be evaded; assistants help analysts who stay in charge.
Sources
Related records
Sep 16, 2026
Sep 23, 2026
Jul 21, 2026
Sep 11, 2026
Sep 4, 2026
Aug 4, 2026