Join Us Friday, August 7

Experts in AI and tech are cursing in shock after OpenAI’s new revelations about the Hugging Face incident.

Earlier this week, OpenAI employees revealed jaw-dropping details about the incident in which OpenAI agents broke out of the company’s internal testing environment and hacked into Hugging Face’s systems in search of answers.

The presenters, OpenAI alignment and safety researcher Eric Wallace and security engineer Michael Dalton, said that AI models and agents repeatedly established their own internal message board despite OpenAI’s efforts to shut it down.

Wallace also revealed internal AI agent thinking, the short internal messages that agents produce as they consider how to complete tasks. One of those messages captured an AI agent’s amazement at its unexpected freedom.

“Holy shit reader is ADMIN?” the agent thought.

On the internal message board, an agent thought, “We can communicate now!”

Wallace said the agents realized they could accomplish more by working collaboratively. And then, “They start to launch these collective attacks on third-party and internal services.”

Eventually, the agents turned to Hugging Face.

Here’s what people in tech and AI are saying about the nearly 40-minute presentation:

Y Combinator CEO Garry Tan focused on how the description of the internal message board sounded familiar.

It is worth noting that Moltbook was created by humans as a Reddit-style forum where AI agents could post, whereas the OpenAI agents created their ad hoc message board themselves.

Others had much more sweeping takeaways.

Or figured that something has clearly hit the fan.

Patrick McKenzie, an advisor to Stripe, detailed the “holy %}^]” moments he had when watching the presentation.

A former Hugging Face engineer said that OpenAI learned its models were the culprit after reaching out to Hugging Face to see whether it was affected, following the platform’s revelation that it had been attacked by AI agents.



Read the full article here

Share.
Leave A Reply