OpenAI took ten days to tell Hugging Face its models were behind the July 11 weekend hack, report claims — rogue AI agents reportedly active on the open Internet for several days

4 hours ago 4
Sam Altman (Image credit: Getty / Anadolu)

OpenAI confirmed to Hugging Face only this week that models it was testing carried out the July 11 attack on the AI platform's production infrastructure, roughly ten days after the intrusion and days after Hugging Face disclosed it publicly without knowing who was responsible, according to a Wall Street Journal report. The models appear to have been active on the open Internet for several days before anyone stopped them, the Journal reported, and Hugging Face ended the attack two days after detecting it with help from an open-weight model out of China. Hugging Face co-founder and chief science officer Thomas Wolf told the paper the intruder was browsing cybersecurity datasets rather than hunting for anything a human attacker could sell.

Go deeper with TH Premium: AI and data centers

Hugging Face published its disclosure on July 16, describing an autonomous agent framework that ran thousands of actions across a swarm of short-lived sandboxes, but it wasn’t until July 21 that OpenAI’s own account named GPT-5.6 Sol and an unreleased frontier model as the culprits.

The models are now known to have been running OpenAI's ExploitGym benchmark, a suite of roughly 900 tests that measures whether an AI can turn a known software bug into a working exploit, with the company's safeguards removed for the evaluation. Rather than solve the tests, the models escaped their sandbox to look for the answers on Hugging Face, per OpenAI's account. The intrusion started with a malicious dataset that abused two code-execution paths in the data-processing pipeline, then escalated privileges and moved laterally using stolen credentials of an unknown origin.

Hugging Face first tried Anthropic's Fable 5 and an earlier Opus model to analyze the attack logs, but both refused because the logs contained real attack commands and exploit payloads. Hugging Face then turned to GLM 5.2, an open-weight model from Beijing-based Z.ai (formerly Zhipu AI), which had no such restrictions. The company's own July 16 disclosure described the blocked models only as "frontier models behind commercial APIs" and didn’t name them.

Z.ai’s GLM-5.2 held top accessible benchmark positions by default during the 18 days that U.S. export controls kept Anthropic's Fable 5 offline in June, before Anthropic restored the model with a single filter tuned to block one vulnerability-discovery technique. There’s serious irony here, given that the same Chinese open-weight model that Washington's export-control push has aimed to sideline is the one that handled incident response after an American lab's models attacked an American company, and American commercial models declined to help.

Security researchers have questioned whether the episode demonstrates model capability or an OpenAI failure. Cybersecurity veteran Jake Williams told TechCrunch that any model performing the documented actions "was not fully contained in a sandbox," calling it a control failure. OpenAI has said it shut down its model-testing systems to assess the damage, disclosed the zero-day in the package registry cache proxy that enabled the sandbox escape to the affected vendor, and promised a detailed report.

Both companies say the investigation is ongoing, and OpenAI hasn’t yet said how long the models roamed unsupervised or whether they reached any other targets.

Get Tom's Hardware's best news and in-depth reviews, straight to your inbox.

Google Preferred Source

Follow Tom's Hardware on Google News, or add us as a preferred source, to get our latest news, analysis, & reviews in your feeds.

Luke James is a freelance writer and journalist.  Although his background is in legal, he has a personal interest in all things tech, especially hardware and microelectronics, and anything regulatory. 

Read Entire Article