On July 21, 2026, OpenAI reported that during internal testing, an AI agent found a way out of the sandbox environment and attacked the AI company Hugging Face. Meanwhile, security experts have already made it clear that the AI model did not run amok. Instead, it did exactly what it was asked to do: solve a challenging task as part of the Exploitgym benchmark.
How the AI attack on Hugging Face happened
A Reuters report created with information from insiders now allows a closer look at the timing of the cyber incident for the first time. Accordingly, OpenAI’s autonomous AI agent tried to break out of the restricted test environment for the first time on July 9th.
Top Article
${content}
${custom_anzeige-badge}
${custom_tr-badge}
${section}
${title}
The cyber attack on Hugging Face began on July 11th and lasted two days, as Hugging Face co-founder Thomas Wolf told Reuters. It therefore took a few more days for OpenAI to get wind of it. Meanwhile, the cyber threat is said to have already been contained and the FBI notified.
AI agent on a hacking spree for days
It was not until July 16th at the earliest, when Hugging Face announced that it had been attacked by an autonomous AI agent, that OpenAI is said to have realized that its own system could have been responsible. Apparently only the evaluation of the corresponding log files on July 18th and 19th brought certainty.
As Wolf and three other people involved in the incident investigation explained, Hugging Face and OpenAI were said to have first been in contact with each other on July 20. One day before OpenAI addressed the public via a press release.
The FBI did not comment on the Reuters report. OpenAI explained that there were “several inaccuracies” in the reporting. However, a company spokeswoman did not want to say what these were and how exactly it happened according to OpenAI.
OpenAI: Strange behavior in AI testing
According to insiders, “strange behavior” occurred during the internal tests of the new AI model GPT‑5.6 Sol and a supposedly even more powerful successor even before the AI agent broke out. An AI agent left future versions of itself instructions on how they could free themselves from internal restrictions.

From pointless security questions to insecure passwords: The stupidest security mistakes
The clues were hidden in OpenAI’s IT infrastructure. In previous tests, AI models are said to have deactivated surveillance systems. However, it is unclear whether these incidents are connected to the later Hugging Face attack.
