OpenAI built GPT-Red to hack its own AI and hid it



OpenAI trained an elite hacker and then locked him in a cage. Their entire job is to break OpenAI’s own AI. The company says it’s too dangerous to allow anyone else near it.

The model is called GPT-Red and OpenAI. detailed it this week. It’s automated red teaming: software that looks for ways to hijack or sabotage other AI systems, so that holes can be patched before they are released. Humans have been doing this work by hand for a long time. It’s OpenAI’s deepest push yet to automate its own. AI securityand GPT-Red does it at machine speed.

OpenAI targeted immediate injectionwhere hidden instructions, buried in an email, web page or file, trick a model into doing something they shouldn’t. He then unleashed the hacker on real targets.

The training dojo

GPT-Red learns by fighting. OpenAI put it in a self-play loop against a squad of defending models. GPT-Red is rewarded for performing an attack; the defenders to defend themselves against one. As defenders realize, GPT-Red must come up with nastier tricks. OpenAI says it invested some of its largest computing runs to date into the model, an amount it considers unprecedented for security work.

It got good. talking to MIT Technology ReviewThe team said GPT-Red found a completely new kind of attack they had never seen before, which they call a “false chain of thought.” Place a false note in a model’s private working memory, tricking them into trusting something that is not true.

“It’s like I told you 1+1=3 and you’ve already verified it,” said OpenAI researcher Chris Choquette-Choo. “The model says, ‘Oh, okay, of course,’ and just spits out 3.”

Hack the vending machine

The tests became physical. In one, GPT-Red attacked Vendy, an AI agent running a real vending machine in the OpenAI office, built by Andon Labs. He changed prices, marked an expensive item down to the 50-cent minimum, and canceled a customer’s order. OpenAI says it has revealed the flaws.

The scores are surprising. Against an older GPT-5, over 90% of GPT-Red’s strongest attacks worked. against the new GPT-5.6less than 23% did. In a repeat test from 2025, GPT-Red handily beat the red human teams, beating 84% of the scenarios down to 13%.

Kept in a cage

OpenAI trained GPT-5.6 against GPT-Red and considers it its most robust model yet against fast injection. But it will not deliver the attacker itself, so its abilities will remain far from the real ones. agent kidnappers. It is not the first laboratory that builds something and decide not to release it.

“It’s not a trivial thing that someone can do easily,” Choquette-Choo said, “just go and train a super attacker using this idea.”

GPT-Network still has blind spots. It is weak in prolonged, back-and-forth attacks, and in hiding instructions within images. And human testers continue to detect things that they miss. “I think the human experience will continue to be very important,” said Jessica Ji, an AI security analyst at Georgetown CSET.

The broader idea is a flywheel: using today’s models to harden tomorrow’s. OpenAI already does this to make its AI smarter. Now it wants security to grow just as quickly. A full article will be delivered later this week.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *