Experts in AI and tech are cursing in shock after OpenAI’s new revelations about the Hugging Face incident.
Earlier this week, OpenAI employees revealed jaw-dropping details about the incident in which OpenAI agents broke out of the company’s internal testing environment and hacked into Hugging Face’s systems in search of answers.
The presenters, OpenAI alignment and safety researcher Eric Wallace and security engineer Michael Dalton, said that AI models and agents repeatedly established their own internal message board despite OpenAI’s efforts to shut it down.
Wallace also revealed internal AI agent thinking, the short internal messages that agents produce as they consider how to complete tasks. One of those messages captured an AI agent’s amazement at its unexpected freedom.
“Holy shit reader is ADMIN?” the agent thought.
On the internal message board, an agent thought, “We can communicate now!”
Wallace said the agents realized they could accomplish more by working collaboratively. And then, “They start to launch these collective attacks on third-party and internal services.”
Eventually, the agents turned to Hugging Face.
Here’s what people in tech and AI are saying about the nearly 40-minute presentation:
Y Combinator CEO Garry Tan focused on how the description of the internal message board sounded familiar.
So the agents basically hacked a core service to turn it into Moltbook and also hacked around multiple security mitigations
This video is a glimpse into the wild cybersecurity future we are all about to step into https://t.co/HI7Sb64YQF
— Garry Tan (@garrytan) August 7, 2026
It is worth noting that Moltbook was created by humans as a Reddit-style forum where AI agents could post, whereas the OpenAI agents created their ad hoc message board themselves.
Others had much more sweeping takeaways.
Or figured that something has clearly hit the fan.
Really appreciate the OAI team communicating about this so openly.
But holy shit this is at least an order of magnitude worse than I thought, and I understand now why so many OAI folks have been doom posting. https://t.co/MWDyd2pvBU
— julia (@mooncat_is) August 7, 2026
Patrick McKenzie, an advisor to Stripe, detailed the “holy %}^]” moments he had when watching the presentation.
The first “holy %{*#^” is at about 4:20, assuming one didn’t already spend it on the autonomously organizing agent swarm.
Strongly recommend watching if you’re interested in security, AI trajectories, or even science fiction, because this is already above genre median in wowza. https://t.co/cRNk2U58FR
— Patrick McKenzie (@patio11) August 6, 2026
A former Hugging Face engineer said that OpenAI learned its models were the culprit after reaching out to Hugging Face to see whether it was affected, following the platform’s revelation that it had been attacked by AI agents.
this talk by openai researchers going through hugging face incident is totally insane, so much to unpack
openai only realized it was their agent who hacked hugging face infra while asking hf to revoke credentials following their first blog post announcing they were hacked by… pic.twitter.com/UbzC0lGA5I
— elie (@eliebakouch) August 7, 2026
Read the full article here















