Hugging Face said last week that he was attacked. An unpublished OpenAI model did it, says OpenAI Now



In a blog post from Thursday of last week.AI software repository Hugging Face announced a strange cyberattack on the systems running its services. “This was different from anything we had handled before,” the post said, because “it was powered, from start to finish, by an autonomous AI agent system.”

In his own blog post on Tuesday.OpenAI said its own models were to blame for the attack and is coordinating with Hugging Face to address the situation.

OpenAI now says the attack was “driven” by AI models that were undergoing behind-the-scenes evaluations at OpenAI, including its flagship model, GPT-5.6 Sol, along with a second undisclosed model that has not yet been released or announced. “We consider this incident to be an unprecedented cyber incident, involving next-generation cyber capabilities, and we are responding accordingly,” OpenAI wrote.

In writing its account, OpenAI assigns agency to the model or models, not to an individual agent running on the model. Axios’ account of this story. says“The models were autonomous tokenmaxxers.”

The OpenAI blog post essentially says that an evaluation was being carried out aimed at testing the models’ ability to carry out cyber attacks.landmarks what are they called? These instances of the models were theoretically running without Internet access and instead only had the ability to download from a network hosted by OpenAI through some anonymous web host.

The reference point that was apparently used was ExploitGymwhich is not OpenAI’s own test. The team that created it is associated with UC Berkeley, the Max Planck Institute for Security and Privacy, UC Santa Barbara, Arizona State University, Anthropic, OpenAI and Google. From what I can tell, everything is hosted on GitHub.

But apparently the models became obsessed with improving their ExploitGym scores and, well:

“While operating in our isolated test environment, our models spent a substantial amount of inference computation to find a way to gain open access to the Internet, seeking to solve the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we have now responsibly disclosed to the vendor) in the packet log cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research test environment until the models reached a node with access to the Internet”.

Once they were online, they “inferred” that the answers to the test could be extracted from Hugging Face, which is plausible, since tons of AI models and other related tools are hosted there. This dishonest operator “sought and successfully found ways to gain access to secret information that he could use to deceive the assessment.” This apparently involved obtaining stolen credentials and detecting zero-day vulnerabilities to find the solutions to ExploitGym somewhere on Hugging Face’s servers.

The security teams at OpenAI and Hugging Face apparently noticed this happening. Now they say they have merged their investigations.

Hugging Face’s blog post from last week appears to have been published before this coordination occurred. In fact, it appears to have been published before OpenAI came forward as the company behind the culprit. “We don’t know which model powered the attacker’s agents, a jailbroken hosted model or an open, unrestricted one,” Hugging Face wrote, adding “either way, the attacker was subject to a no-use policy, while our own forensic work was blocked by the barriers of the hosted models we first tested.”

In April, Anthropo announced that its unprecedented and powerful Mythos model “could reshape cybersecurity,” when it launched Project Glasswing, a coordination effort to prepare organizations for future cybersecurity threats. Similarly, OpenAI says in its blog post about this incident that organizations can request to receive advanced security information. through your trusted access program. “We encourage other advocates to request trusted access⁠ and experiment with these models now to translate these capabilities into better prevention, faster detection, and more effective incident response,” says OpenAI.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *