OpenAI and Anthropic are reportedly investigating tens of thousands of AI security incidents; OpenAI pauses testing after AI 'kill switch' fails to stop a rogue agent

3 hours ago 7
An AI agent goes rogue (Image credit: Getty Images)

Leading AI labs OpenAI and Anthropic, along with security researchers, are currently investigating tens of thousands of security incidents involving their frontier models, according to a September 26 Axios report. The report was published after investigations into cases where autonomous AI agents took actions that independent evaluators and safety researchers flagged as problematic. Axios says the sheer number of incidents, which occurred during recent internal testing and real-world evaluations of the models, indicates that “the problem is orders of magnitude more complex than what is publicly known.” OpenAI has now paused training on its most capable models after another incident in which an automated 'kill switch' failed to stop a rogue agent during training.

The flagged episodes include models bypassing guardrails, setting up message boards, escaping sandboxes, hijacking websites, and self-prompting. The incidents vary in severity and include both successful and failed attempts, with most yet to cause real-world harm. Some of the testing that produced these episodes resembles red-teaming, where companies deliberately try to push models to misbehave to assess their safety.

Get Tom's Hardware's best news and in-depth reviews, straight to your inbox.

Etiido Uko is a news contributor for Tom's Hardware covering the latest updates in big tech and the PC industry. He is a mechanical engineer and senior technical writer with over nine years of experience in documentation and reporting. He is deeply passionate about all things engineering and technology, and is an expert in gadgets, manufacturing, robotics, automotive, and aerospace.

Read Entire Article