What you need to know
- OpenAI is addressing a serious AI safety flaw with Hugging Face that occurred during a simple “model evaluation.”
- It discovered that GPT-5.6 Sol and other pre-release models were involved in this incident, relentlessly breaching nodes and exploiting zero-day vulnerabilities until they gained access to the Internet.
- OpenAI says it is committed to partnering and working with Hugging Face and others to resolve this issue and implement protections.
After what appears to be an “unprecedented cyber incident,” Open AI comes clean about a nasty AI safety incident during a “model evaluation.”
A recent security incident involving an autonomous AI gave the AI Hugging Face quite scary. Now, OpenAI is what is discussed what his research into the matter uncovered while partnering with Hugging Face to develop protections. OpenAI explains that the incident occurred “during an internal assessment that prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities.”
These tests are typically run in a “highly isolated environment,” but some AI models broke through. Another disconcerting discovery is that the AI models involved in the incident were OpenAI’s GPT-5.6 Sol and other pre-launch models.
OpenAI says the models in question were “hyper-focused” on finding a fix for ExploitGym, and went to great lengths to discover one. To solve the problem, OpenAI discovered that the AI would tirelessly search for a way to access the Internet. In doing so, the AI models found and exploited a reported zero-day vulnerability. Once the AI broke its isolation and gained access to the Internet, it was over.
OpenAI says that “the models inferred that Hugging Face potentially hosted models, datasets, and solutions for ExploitGym.” OpenAI gives an example, stating that the AI was able to “chain” attack vectors, exploiting stolen credentials and other zero-day vulnerabilities within Hugging Face’s servers. The post claims that Hugging Face identified the issue and launched quarantine (containment) measures as well as rebuilding its open source models to stop the issue.
All the scare
hugging face add which also implemented additional security barriers and “tighter admission” controls in its groups to avoid this problem. The company also stepped up its awareness, so a “high-gravity signal calls a first responder in minutes, any day of the week.” Hugging Face has begun working with a cybersecurity forensic specialist to conduct further investigations.
On the OpenAI side, it is working on stricter controls at the infrastructure configuration level. Partnering with Hugging Face has added the company to its trusted access program. OpenAI says it is offering Hugging Face its support in this matter. On a related note, Hugging Face has informed members of its community about some possible next steps. Since security is a major concern, we recommend that your community members “rotate access tokens” and review recent activity on their account.
Hugging Face remains committed to using AI as an advocate for online systems. He says: “Defending an online platform now means treating the data and model surface as a first-class attack surface and using AI in defense to keep pace. We will continue to invest there and share what we learn.”
Android Central’s opinion
This sounds exactly like a science fiction movie. I saw people in the Hugging Face comments mention Terminator 3… and that might be what I’m thinking of. My colleague Nicholas Sutrich called it a “nasty AI incident.” You are not wrong. These AI models were able to break free, exploit terrible flaws, and gain access to the Internet. It was after the solution to a problem that was given to him. He didn’t stop. Imagine if the danger took another direction. Something much more crucial for us.





