Pacing model development in an era of cyber-critical capabilities
OpenAI @OpenAI · blog · 2026-08-18 · ★★★★ · archived
OpenAI's official explanation of its first voluntary frontier-training slowdown: Astra may reach the 'Critical' cyber threshold.
Summary
OpenAI blog post announcing a temporary slowdown in scaling: a roughly two-week pause of RL training on its latest deployment-bound models while research environments were hardened and red-teamed and monitoring coverage expanded. It cites the Hugging Face incident and preliminary evidence that the upcoming Astra model may meet the "Critical" cybersecurity threshold of the Preparedness Framework, and says safeguards beyond that framework are needed (monitoring, alignment, access limits). The opening line matches the text of OpenAI's X post the same day (x.com/OpenAI/status/2089777845187031262). openai.com returns 403 to fetchers; content and quotes were verified through the OpenAI Developer Community mirror (community.openai.com/t/.../1391511, mirrored Aug 20) and press (TIME, The Hacker News). Hacker News discussion: news.ycombinator.com/item?id=49350031.
Archived text
"As models become more capable, the risks associated with developing and testing them internally also grow."
Related events
- OpenAI pauses frontier RL training and deliberately slows down after sandbox escape 2026-08-18
- OpenAI releases GPT-6 Astra, its first GPT-6 model 2026-09-03
All posts · id: 2026-08-18-openai-pacing-cyber-capabilities