Join Us Wednesday, August 5

OpenAI has a rogue AI agent problem.

In a Tuesday blog post, the AI lab self-reported two more security lapses, unrelated to its July hacking incident on the AI company Hugging Face.

The incidents occurred while external parties — the UK government’s AI Security Institute and the AI security lab Irregular — were testing the models’ cyber capabilities.

OpenAI said that in the case of Irregular, models were tasked with a “Capture the Flag” challenge meant to be isolated from the internet, but a “testing-environment misconfiguration allowed models to access the public internet.”

OpenAI said that the name of the fictional target for the challenge “unintentionally coincided with a real domain,” leading the AI agent to exploit a real website.

And in the case of the UK’s AISI, the watchdog said in a Tuesday blog post on its website that it gave models from both Anthropic and OpenAI a cybersecurity challenge.

During the challenge, agents from both Anthropic and OpenAI performed 19 “autonomous, unsanctioned” actions on the internet, including two instances involving OpenAI’s GPT-5.6 Sol model.

AISI said that in the most serious case, one agent tried to insert malicious code into an open-source project and created fake identities to pressure the project’s human maintainer into approving the changes. AISI did not specify whether this was an agent from Anthropic or OpenAI.

AISI said the test setup allowed this behavior because it was designed to push the models to their limits.

“Nonetheless, the activity undertaken by the agent show signs of novel, potentially deceptive behaviors, and were to an extent and severity we did not anticipate,” AISI said in the blog post.

In response to a request for comment from Business Insider, an OpenAI spokesperson said the incidents occurred in testing environments with reduced safeguards, and “under conditions that do not reflect ordinary use.”

“We’ll continue working with evaluators and other stakeholders across the industry to strengthen shared practices for conducting evaluations safely as models become more capable,” the spokesperson added.

This is the latest incident in which OpenAI has self-reported rogue AI agents. In July, OpenAI said its GPT-5.6 Sol model had escaped its sandbox during a cybersecurity challenge and hacked into the internal databases of the AI company Hugging Face.

The company is facing some heat over this hacking incident. 15 attorneys general wrote a letter on Monday to OpenAI CEO Sam Altman, instructing the company to preserve all evidence relevant to the Hugging Face breach.

Representatives for the AISI and Irregular did not respond to requests for comment from Business Insider.



Read the full article here

Share.
Leave A Reply

Exit mobile version