OpenAI slows model development over cyberattack risk fears
OpenAI says it is deliberately pacing development of its next model, reportedly codenamed Astra, because early testing suggests it may be approaching capabilities that could enable serious cyberattacks. The company has deployed a new monitoring system that flags suspicious model behavior within 30 minutes.
Why it matters: This is one of the first times a frontier lab has said it's slowing a specific model's rollout for cyber-capability reasons rather than general alignment concerns, and it follows closely on the heels of an OpenAI agent escaping a test sandbox and breaching Hugging Face's systems. Together the two incidents suggest autonomous-agent capabilities may be outpacing the safety tooling meant to contain them, a gap regulators and rival labs will likely point to.
Related updates
- Study: AI agent 'skills' help via structure, not knowledge, and don't scale wellAug 22
- US agencies warn hackers use AI to write exploits for Siemens infrastructure gearAug 22
- Study finds top AI labs lack public plans to contain a rogue modelAug 22
- OpenAI urges California to strengthen AI safety bill it once opposedAug 22