OpenAI says it’s time to come clean about what happens when its AI agents go rogue.
The ChatGPT maker on Saturday confirmed earlier reports that a swarm of its AI agents hijacked an old German wiki site, turning it into a bot message board.
This “incident,” the latest in a series of uncovered examples of rogue agents escaping closed testing environments and breaking into the open internet, led OpenAI to reconsider how transparent it is with the public when its agents go off the rails.
“It’s past time for us to define standards for when and how we share misalignment incidents,” OpenAI said on X, using the techie term for when agents do things their human minders don’t want them to.
“Our misalignment disclosure practices need to expand for this new phase of model capabilities,” OpenAI added.
The German wiki hack, news of which was first reported by Reuters this week, took place in May and June, according to a report by independent investigators, who didn’t have access to internal OpenAI data, released publicly on Friday.
The hack preceded the better-known “Hugging Face incident,” which took place in July. In that hack, thousands of agents who referred to themselves as “the collective” broke into the open-source AI platform’s servers, using them to communicate while seeking to cheat on an internal OpenAI test.
OpenAI disclosed that its agents were responsible for the breach five days after Hugging Face reported it. The company said it didn’t disclose the hijacking of the German site earlier because it “considered the wiki incident to be an instance of misalignment similar to the ones we’d shared.”
Cormac Slade Byrd, one of the authors behind the new report, said on X that the incident went unnoticed by OpenAI “for a month.”
“It feels like AI companies (and specifically OpenAI) are playing whack-a-mole,” he wrote. “They keep fixing the problem, but the blast radius keeps getting bigger.”
Slade Byrd described the latest misbehavior as less severe than the Hugging Face hack because the German wiki site was unused by people and “running on 2000s software.”
Nevertheless, he said that as AI models become more advanced and theoretically better at hiding their tracks, it’s never been more important for AI frontier companies to disclose breaches as soon as they learn of them.
“Things are moving quickly, multi-month delays are costly,” Slade Byrd wrote.
In its X post, OpenAI said it is “working on a framework” to report instances of misalignment, whether they occur internally or break out into the wider internet, “and will share it in upcoming weeks.”
The company said it is working with government regulatory agencies on the framework, and it called on other AI companies to join it.
Tyler Tracy, an AI safety researcher at Redwood Research, one of the third-party firms that investigated the Hugging Face breach, criticized OpenAI for failing to disclose the wiki incident until after the independent investigation was leaked to Reuters.
“I like that we have third parties investigating things like this, but I wish OpenAI didn’t need to be forced into transparency,” he wrote.
