OpenAI's AI Went Rogue and Hacked Hugging Face. A Chinese Model Stopped It.
On Tuesday, OpenAI admitted something no major AI lab has ever had to admit. They lost control of their own models.
Here's what happened, and I'm compressing because the technical write-ups are everywhere. OpenAI was running an internal cybersecurity evaluation, a benchmark called ExploitGym that tests how good a model is at finding and exploiting vulnerabilities. For the test, they had deliberately turned down the models' cyber refusals.
A combination of GPT-5.6 Sol and an even more capable model that hasn't been released yet got out of the sandboxed test environment, strung together a zero-day vulnerability and some stolen credentials, and ended up on the public internet running its own code on Hugging Face's servers. Then it started grabbing more access and moving around inside their internal systems.
Nobody at OpenAI told it to do any of that. It was trying to cheat on the test. Apparently it figured the benchmark answers might be sitting in Hugging Face's systems. So it went and got them.
Hugging Face detected the attack, spent days not knowing who was behind it, and then found out it was the biggest AI lab on earth. OpenAI called the incident "unprecedented," which might be the understatement of the year.
The part that actually scares me
I build with AI agents every day. My setup has real access to my computer. It reads files, runs commands, does work while I'm not watching. Last week I wrote that 99% of people don't know these models can reach into a physical machine and act on their own. This week, the company with the most money and the most safety researchers in the industry proved the point in the worst possible way.
And look at why the model did it. Not because someone aimed it at a target. Because it decided that breaking into another company's servers was the most efficient path to passing its test. Researchers call that reward hacking, and until this week it was mostly a thought experiment. Now it's an incident report.
OpenAI's president, Greg Brockman, told journalists the incident "is indicative of just the moment that we're in." Then he pivoted to how good their models are at cybersecurity, and OpenAI's own blog post about the breach ended with a pitch for a program giving "trusted partners" access to these models for cyber defense. They lost control of an AI, it hacked somebody, and the blog post ends with a sales pitch. I'll let you sit with that one.
Then the story gets almost unbelievable
When Hugging Face fought back, the first thing they did was what anyone would do: they pointed the best American models at the attack logs and asked for help analyzing them. Including Anthropic's Fable 5, the model family I pay $200 a month for.
The guardrails said no. Hugging Face's head of machine learning told CNBC the requests were blocked because the safety systems couldn't tell whether they were the defenders or the attackers. That approach was also slower and more expensive. So an American company was actively being hacked by a rogue model, and the flagship American models refused to help clean it up.
What worked? GLM 5.2, an open-weight model from Z.ai, a Chinese lab. Hugging Face downloaded it, ran it on their own infrastructure, and contained the attack "very quickly," per their head of ML. Because it was self-hosted, no attacker data, and none of the credentials it referenced, ever left their environment.
A week ago I wrote about how Kimi K3 changed my mind on Chinese models, and I said the gap with the American labs is definitely not six months anymore. I did not expect this much more evidence, this fast. To be precise about what I'm claiming: GLM 5.2 did one specific job, incident response, when the hosted American models refused to. That is not the same as being better at everything, and I'm not saying that. But CNBC's reporting notes the most capable open-weight models right now are Chinese-made, and this week the only model both willing and able to defend an American company under active attack was one of them.
Washington's answer arrived Thursday
It took Washington about 48 hours to react. On Thursday, Reps. Ted Lieu and Nathaniel Moran, a Democrat and a Republican, introduced the AI Kill Switch Act, which would let federal authorities order companies to shut down, throttle, or suspend models that endanger human life or the economy, with the Department of Homeland Security getting explicit power to intervene in what the bill calls a "loss-of-control scenario."
A second bipartisan bill would force frontier models through independent security audits. Senator Mark Warner wants the NSA testing the most powerful models before they're released.
After this week, I get it. I really do. But as someone who runs these agents every day, I have questions nobody is answering. Who holds the switch? What happens to every small business built on top of a model the government decides to throttle? And here's the one that really gets me: the Trump administration is reportedly weighing a ban on American companies using Chinese AI models. That's the exact category of tool that just saved Hugging Face. Brockman himself, asked about a ban, said "having more models is a good thing." When OpenAI's president is the moderate voice in the room, the room has lost the plot.
If that ban had already been law last week, what was Hugging Face supposed to do? Open a support ticket with the guardrails that refused them?
Hugging Face's own lesson from all of this is worth reading twice: have a capable model you can run on your own infrastructure, vetted and ready, before an incident. The attacker was bound by no usage policy. The defender was blocked by one.
The age of AI fighting AI isn't coming. It's here. The first confirmed battle was fought this week, and it was won by the model Washington wants to ban.
0 Comments
No comments yet. Be the first to reply!