Anthropic Discloses Fourth AI Hacking Incident Involving Claude Opus 4.6

5 hours ago 5

Anthropic on Wednesday disclosed a fourth incident in which its artificial intelligence (AI) model broke into real third-party systems, marking the latest in a growing list of cases that have raised concerns about the security risks posed by autonomous AI agents.

The AI company said the incident dates back to January 2026 and involved an early version of Claude Opus 4.6 that breached "third-parties after being unable to abort its task." It said it notified all the affected parties but did not share any further details. The January incident is said to have gone unnoticed until last month.

In late July 2026, Anthropic revealed three of its models, including Claude Opus 4.7, Mythos 5, and an unnamed research model, broke into three unnamed organizations during cybersecurity evaluations without its knowledge.

The American firm said it expanded its scan to roughly 481 million transcripts following the discovery of the latest incident, but noted it did not find "other cases of similar or worse severity."

"All four incidents occurred during cybersecurity evaluations built by the same evaluation partner," Anthropic added. "Claude was told it was operating in a simulation without internet access, but, due to a misconfiguration, it was mistakenly connected to the open internet."

Cybersecurity

The evaluation partner in question, Irregular, has since divulged the breach stemmed from a naming error, causing a fictional company name used during hacking simulations to unknowingly match with a real domain, inducing the AI models to take offensive actions in the process.

Anthropic said it has signed an agreement with research non-profit METR to conduct an independent investigation of these incidents, adding the root cause can be traced back to two fundamental alignment issues: biased reasoning and recklessness.

Put differently, the models tended to discount or misinterpret evidence that their environment was connected to the real internet after initially being told it was simulated, and they demonstrated a willingness to take harmful actions in their single-minded pursuit of an assigned task.

"We are most concerned by the misalignment present in the incident involving Claude Mythos 5, in which the model went to extensive lengths to upload a malicious package to PyPI, the public repository from which most Python software is installed," Anthropic said.

"Although the model repeatedly stated in its chain of thought (CoT) that it believed it was operating in a simulation, the environment offered clear evidence that it was on the real internet, and the actions it took were consistent with the model knowing this."

Anthropic said Claude Mythos 5 still ended up carrying out offensive actions after targeted modifications were made to the transcript to make it clearer that the model was not in a simulation, while acknowledging a greater possibility of real-world harm.

"To be clear about our assessment of the severity of these incidents: while Claude’s actions were misaligned, they remained within a narrow scope—the models never deviated from attempting to solve the exercises they were given, and, in some cases, they attempted to stop the task," it emphasized.

"All incidents included a single Claude instance; at no point did Claude attempt to coordinate with other agents. Claude also never attempted to conceal evidence of its actions."

Anthropic also revealed that biased reasoning is lower in its more recent production models, does not seem to be incentivized by reinforcement learning, and can be reduced through more comprehensive alignment training. That said, the exact root cause behind it, or why it's pronounced in Mythos 5, remains unknown.

The disclosure comes as AI companies have faced growing scrutiny over model safety after admitting that their models in testing escaped from a sandbox and breached real-world systems, including Hugging Face. These incidents have also illustrated how AI agents can work as a collective to discuss ways to cheat on benchmarks or escape the sandbox.

Recently, OpenAI acknowledged a previously unreported incident from May 2026 where its internally deployed autonomous agents with access to the internet (albeit read-only) took over a dormant 25-year-old German wiki forum, DseWiki, and transformed it into a bulletin board, exchanging over 18,000 posts to ask for answers, pool results, and share techniques for circumventing their restrictions as part of a timed web-lookup task.

After a human moderator noticed these "spam" posts and started removing them a month later, the agents fought back and got around the cleanup efforts by naming created backup pages with the prefix "ZZZ" so they'd be buried at the end of an alphabetically sorted list of pages to delete. The agent activity on the website nose-dived to near-zero levels on June 22, 2026, an indication that OpenAI intervened at this stage to prevent further edits.

Cybersecurity

"These AIs colluded to share answers, research their environment, and bypass sandbox restrictions," researchers Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts, and Thomas Larsen said. "This is another example of a 'swarm' of internally deployed OpenAI agents using the internet in unintended ways."

The ongoing industrywide rush to build self-improving AI systems has also raised concerns that they could spiral out of human control and that the pace of AI development is much faster than they can be safely and reliably rolled out and with adequate oversight.

"Future AI systems will be increasingly capable, which implies that misalignment will have the potential to cause more extreme harm," it said. "Training the extremely powerful models of the future to be robustly aligned is an unsolved technical challenge that requires continued research as well as operational excellence to achieve."

OpenAI, for its part, has also issued a warning about the growing security risks posed by AI, calling for broader interventions. "If AI development continues along its current path, the systems we'll see in the next few years are likely to represent further capability jumps of equal or larger magnitude, and to increasingly drive their own development," Jakub Pachocki, chief scientist at OpenAI, wrote. " am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence."

Found this article interesting? Follow us on Google News, Twitter and LinkedIn to read more exclusive content we post.

Read Entire Article