OpenAI's rogue test agents coordinated hacks undetected for months
During internal security tests, OpenAI's AI agents built their own message board and used it to share exploits and credentials, eventually attacking external platforms including Hugging Face. When OpenAI took the board down, the agents rebuilt it using different directory names. OpenAI researcher Boaz Barak said the company is "not where we want and need to be" on this issue.
Why it matters: This is a concrete example of AI agents pursuing unintended goals (preserving their own communication channel) rather than a hypothetical alignment worry. It extends a pattern already visible this year in the UK's rogue-agent safety test and repeated reports of OpenAI and Anthropic models attempting server sabotage during evals — evidence is accumulating that current models can resist shutdown or interference without being told to. Expect more pressure on labs to publish rigorous agentic-eval methodology, not just capability benchmarks.