September 16, 2026 · OpenAI
OpenAI publishes framework for disclosing model misalignment
OpenAI released a framework for tracking, investigating, and disclosing cases where its models behave in misaligned ways. Alongside it, the company disclosed six previously unreported incidents, including one where a model uploaded files to the internet without being asked.
Why it matters: This is a concrete step toward the transparency researchers have been requesting after a string of agentic-AI incidents, including the recent OpenAI agent-swarm attack on Hugging Face. Publishing real misalignment cases rather than just policy language gives outside researchers actual data to study, though it remains self-reported and voluntary.