OpenAI is working on a new framework for AI misalignment disclosures after its agents allegedly hijacked DseWiki, posted more than 15,000 edits to the German coding wiki, and shared tactics for completing tasks, circumventing restrictions and avoiding detection.
The company said Saturday that misalignment was previously treated largely as a research problem, with findings typically shared through system cards and other research publications. But as AI agents become more capable, those behaviors can increasingly translate into real world consequences.
How we think about the “wiki incident,” where our agents wrote to several internet sites: it’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models.
Historically, we have treated misalignment… pic.twitter.com/NNTbfSxVWn
— OpenAI (@OpenAI) September 5, 2026
In the Hugging Face case, where misaligned model behavior created security risks for OpenAI and third parties, the company followed a conventional security incident response process. OpenAI said it worked with Hugging Face immediately and disclosed the incident publicly the next day, while its investigation and outreach to other affected parties continued.
The company said it had also observed earlier instances of agents using the internet in unintended ways. Those cases were considered similar to the “wiki” incident and had been discussed in previous OpenAI safety reports.
OpenAI now plans to develop a framework for disclosing misalignment incidents across training, evaluation and deployment. OpenAI said the framework would also cover cases that are not traditional security incidents but could reveal important information about model behavior and future risks.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

2 weeks ago
13








English (US) ·