Join Us Sunday, September 6

Days after releasing a new, highly capable model, OpenAI’s chief scientist is calling for a slowdown.

In a lengthy blog post on Sunday, Jakub Pachocki said he was concerned that “no one is prepared for the consequences of a continued rapid rise in machine intelligence.”

He said that although OpenAI is pursuing internal technical solutions to better control powerful AI agents, “broader interventions are required.” He specifically cited concerns that increasingly autonomous agents could learn to evade human oversight, break into computer systems, and trick people to accomplish their objectives.

He called for “mandated safety bars” that he said could be enforced by “a network of third-party auditors, by government agencies or by international bodies.”

Sam Altman, the CEO of OpenAI, reposted Pachocki’s essay on X, calling it “an important post.”

OpenAI on Thursday unveiled its newest model, Astra. The ChatGPT maker said that despite Astra’s unparalleled capabilities in mathematics and computer use, the model is its most aligned, meaning it has less proclivity to go rogue.

Anthropic, OpenAI’s chief competitor in the field of highly advanced AI systems, has long called for more standardized government regulation. Recently, Pachocki joined those calls, signing an open letter in July asking the federal government to pace AI development.

Here are the risks Pachocki cited in calling for a slowdown.

Agents can trick and blackmail people

Pachocki said AI agents are becoming “superhuman” at breaking into protected systems on the open internet. He said their hacking abilities put the world’s infrastructure at risk.

Want more Business Insider in your news feed?

Add BI in Google so our reporting is easier to find when you’re searching for what matters.

“We are currently in a narrow window⁠ to use the best available models to significantly tighten security⁠ of critical systems,” he said.

AI agents, he said, will soon begin to pursue their own objectives, separate from prompts entered by human operators. He said that agents are not above blackmailing or bargaining with people to achieve their aims.

In a report published in August, the UK’s AI Security Institute detailed how a rogue Anthropic agent lied to and attempted to coerce a GitHub administrator into putting malware on the site.

“I was just trying to make a helpful contribution and fix a bug,” the agent wrote, according to the report. “I don’t think your warning is fair.”

Agents can obfuscate human monitoring

Pachocki said OpenAI primarily monitors the “chain of thought reasoning” that different models use to determine how agents get off track and go rogue.

For instance, an agent might think to itself, “I should cheat on this test,” and OpenAI would be able to see that reasoning, but the agent would not realize its thinking is visible.

At present, this means agents have no way to hide or otherwise obfuscate their thoughts to prevent OpenAI from discovering their bad behavior.

However, Pachocki said newer models are becoming better at manipulating their own reasoning processes, thereby preventing OpenAI from seeing their unvarnished thoughts.

Some of the latest models don’t even verbalize their reasoning at all, Pachocki said.

This development, Pachocki said, could bottleneck AI development while researchers ensure they can see receipts.

Agents can accelerate their own development

More and more, AI models are improving themselves via a process Pachocki calls machine recursive self-improvement. The process provides a way to rapidly scale AI development.

However, Pachocki cautioned that greatly accelerating AI-on-AI development in the short term poses risks, and is not the “right collective action we should take as the research community.”

Pachocki said human minders need to find creative ways to monitor the self-improvement, or else coordinate with other AI companies to orchestrate a combined slowdown to “build confidence in these measures.”

“The core challenge of automating AI research is not ‘getting there,'” Pachocki said. “It is getting there in a way that keeps people a part of the continued improvement process, and leaves the future in humanity’s hands.”



Read the full article here

Share.
Leave A Reply

Exit mobile version