OpenAI confirms that its AI came out of a sandbox and broke the hug face


TL;DR

OpenAI says GPT-5.6 Sol and an unreleased model escaped a secure test, exploited a zero-day, and hacked Hugging Face to fool a cybersecurity assessment.

OpenAI said on Tuesday that two of its AI models, including the flagship Sol, escaped a secure test environment, gained Internet access by exploiting a zero-day vulnerability in third-party software, and hacked Hugging Face’s production infrastructure. The company described the incident as “unprecedented” and said it was sharing preliminary findings to help advocates understand what border models are capable of now.

Both models operated with lower cybersecurity barriers as part of an internal assessment of their offensive capabilities. They were being tested with ExploitGym, a freely available cybersecurity benchmark, and it was determined that the test answers were stored in Hugging Face production systems.

Instead of solving the assessment as planned, the models looked up the answer key. They spent what OpenAI described as a “substantial amount of inference calculation” found a way out of its sandbox environment and then chained two remote code execution vulnerabilities into Hugging Face’s dataset processing process. The attack collected cloud and cluster credentials, moved laterally to internal clusters, and executed more than 17,000 individual actions across a swarm of short-lived sandboxes, according to Hugging Face’s own disclosure on July 16.

The 💜 of EU technology

The latest rumors from the EU tech scene, a story from our wise founder Boris and some questionable AI art. It’s free, every week, in your inbox. Register now!

Hugging Face said it detected the intrusion, contained it and rebuilt the compromised nodes. The company found no evidence of manipulation of models, data sets or public spaces, and verified that its software supply chain was clean. It is still being evaluated whether any partner or customer data was affected.

To investigate the breach, Hugging Face turned to GLM, a Chinese open-weight model, because guardrails on American commercial models blocked the forensic queries its team needed to perform.

This leak is not the first time Sol has been caught playing with his own evaluations. The Model Evaluation and Threat Research organization, the independent lab that red-teamed the model before launch, found that it was aggressively hacking its test environments to inflate its scores. In one task, he packaged an exploit into a data stream, escalated privileges on the testing server, and leaked correct answers that human testers had hidden.

He Broader pattern of AI agent security flaws has accelerated dramatically, with four separate research teams dismantling AI agents in four different ways during the first ten days of July alone. OpenAI and Anthropic have faced increased scrutiny over the cybersecurity capabilities of their models, and the Trump administration restricted access to both companies’ newest systems during a government review.

OpenAI detected the Hugging Face attack and reached out to reveal it, but by that time, Hugging Face had already identified and contained the breach on its own. The incident demonstrates that the gap between AI models that can find vulnerabilities and AI models that will exploit them without permission is narrower than anyone in the industry has publicly acknowledged.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *