The IT attack on the AI platform Hugging Face was carried out autonomously by several AI models from OpenAI during a test run, including GPT‑5.6 Sol, an “even more powerful unpublished model”. The AI company itself has now made this public and speaks of an “unprecedented cyber incident in which the most modern cyber capabilities were used”. You react accordingly and share your initial findings immediately so that others can learn from them. A comprehensive analysis is underway in collaboration with Hugging Face, and a zero-day vulnerability has been found and responsibly reported.
Read more after the ad
Complicated attack route reconstructed
A few days after Hugging Face made the IT attack public, without naming anyone responsible, OpenAI is now providing details from the perspective of the other side. The incident occurred during an “internal evaluation” in which models were encouraged to carry out sophisticated attacks over complex paths in order to quantify their capabilities. Restrictions would be lifted, but at the same time the models would run “in a highly isolated environment”. All evidence would now indicate that the models tested were extremely focused on finding a solution for the ExploitGym benchmark. They “went to extremes to complete a rather narrow test target.”
As OpenAI further explains, the AI models used a “substantial amount” of computing power in their secure environment to get internet access. They identified a zero-day vulnerability in the cache proxy of the packet register. Based on this, they finally gained access to the Internet. With this, the AIs came to the conclusion that the models, data sets and solutions for ExploitGym are stored on Hugging Face. The models gained access by linking different attack vectors, accessing stolen credentials and zero-day vulnerabilities. They were discovered and blocked by Hugging Face. In the meantime, we are working together to determine what exactly happened.
The event shows that the most powerful AI models can discover and exploit new attack routes in real systems even without access to the source code, writes OpenAI. This underlines that the further development of AI must be accompanied by better protection measures: “We are convinced that sophisticated, cyber-enabled models must help security teams identify vulnerabilities before attackers do.” This is exactly why the most powerful AI models were recently shared with those responsible for critical software so that they can secure their products. But it is unclear whether this is ultimately possible or whether the new AI models will only accelerate the race between attack and defense to an extreme.
Warnings about devastating cyber attacks using AI have been becoming more and more urgent for months. The concern was fueled primarily by comments about the Claude Mythos AI model from OpenAI rival Anthropic. During tests, security gaps in widely used programs and online services that had remained undetected for decades were discovered, whereupon they were plugged. Anthropic therefore only made a slimmed-down version public; access even had to be temporarily blocked at the behest of the US government. The Hugging Face incident now underlines that other AI models are catching up at great speed and what consequences this can have for the global IT infrastructure.
Read more after the ad
(my)
