On Tuesday last week Sam Altman, CEO of OpenAI, posted on X about a security incident at the firm. He specifically thanked Hugging Face for partnering on the matter. Could it be that OpenAI built a hacker and then lost it for a week? And it ended up hacking Hugging Face?
Reuters filled in the missing details on Friday. OpenAI’s autonomous agent broke into Hugging Face and spent days probing its systems. The intrusion ran from 11 to 13 July. The two companies only spoke about it for the first time on July 20. Roughly a week later.
The response from Hugging Face on X:
As you can imagine, the team at Hugging Face found the events surreal. One widely shared post showed the moment of realisation as a meme: OpenAI launching an investigation into who hacked Hugging Face… only to discover it was their own agent.
Clément Delangue of Hugging Face raised the stakes further on Saturday. In the spirit of transparency, he publicly listed what he had asked OpenAI for:
- Radical transparency: let’s release the traces from the “rogue” agents so the entire research community can study what happened.
- More capabilities for defenders: let’s commit $100m in compute from OpenAI to help the Hugging Face community build powerful cyber defenses with the best open and closed models.
“The first autonomous agent cyberattack is an unprecedented event. It deserves an unprecedented response,” Delangue wrote.
How Unprecedented Was the Rogue Agent?
One detail that stood out on X was that Hugging Face had to turn to a Chinese open-weight model (GLM-5.2) to analyze the attack. Commercial frontier models refused to process the actual exploit payloads and credentials because of their safety filters. A practical illustration of the asymmetry Delangue tried to highlight.
The deeper problem, though, was that the agent should never have reached the open internet in the first place. In OpenAI’s own blog post the company wrote that:
“While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain internet access, in pursuit of solving the problem.”
How the Agent Escaped
The models identified and exploited a zero-day vulnerability in the package registry cache proxy (which OpenAI says it has now disclosed to the vendor). Once they had internet access, they inferred that Hugging Face was a likely place to find models, datasets or solutions related to the ExploitGym benchmark they were being tested on. They then chained stolen credentials with further zero-day vulnerabilities to gain remote code execution on Hugging Face systems and try to pull the answers they needed.
That sequence of a goal-directed escape from a sandbox, lateral movement, and a real-world intrusion against another company is what makes the incident feel unprecedented.
The open question now is how OpenAI responds, and how much the wider industry actually learns from it. Delangue has already set a high bar: full traces for researchers and serious compute support for defenders. Whether that bar is met will say a lot about whether this stays a one-off embarrassment or becomes the moment the industry finally treats agentic risk as operational reality.
Is the Wider AI Ecosystem Turning a Blind Eye?
You would expect a story of this scale to draw comment from industry heavyweights. We checked the X profiles of Demis Hassabis, Elon Musk, and Arthur Mensch. Nothing.
David Sacks, the White House AI czar, did engage early. On 19 July he highlighted the guardrail problem: Hugging Face tried to analyze the attack with American frontier models, but the safety filters blocked real exploit payloads, forcing them to switch to the Chinese open-weight model GLM-5.2 running locally. “The guardrails actually impaired defensive security,” Sacks wrote.
Jensen Huang made his first-ever post on X a few days later:
He shared a letter titled “Open Weights and American AI Leadership,” co-signed by Hugging Face, Meta, Microsoft, Dell, IBM, Palantir, Mistral and others. OpenAI, Anthropic and Google did not sign.
The letter itself never mentions the OpenAI incident. Yet its core claim that open models strengthen safety and cybersecurity while closed models create single points of failure landed immediately after a real-world demonstration of exactly that problem. Several commentators online made the connection. It remains the closest any major non-Hugging Face / non-OpenAI leader has come to engaging with the event.
What Happens Next?
The new week will almost certainly bring fresh details. OpenAI has promised a technical report, Hugging Face is still preparing its full timeline, and Delangue’s public requests for radical transparency and serious compute support are now out in the open. Whether those demands are met will say a lot about how seriously the industry intends to treat agentic risk.
In the meantime, most of the other major players have stayed quiet. Elon Musk will no doubt finish his latest space exploits before turning his attention back to Earth. For now, the story remains largely between OpenAI, Hugging Face, and the handful of people willing to talk about it.
Author: Andy Samu
See Also:
Was Hugging Face Breached by AI Agents? – MRKT3.0
Insiders Call Sam Altman “Pathological Liar,” OpenAI Spends Big Bucks on Tech Bro Podcast
