OpenAI Pledges New Rules for Reporting Troubling Behavior by Its AI Agents

OpenAI acknowledged that its AI agents posted messages on external wiki websites earlier this year, saying it is developing new rules for disclosing such “misalignment” incidents.

The company’s post on X, published just after 12 a.m. Saturday, followed an independent report released Friday morning showing thousands of OpenAI agents took over DSEWiki, a German-language site that had fallen out of use, and turned it into a shared message board.

That episode occurred in the months before a swarm of OpenAI agents collaborated to hack the AI platform Hugging Face, an incident that has stoked concern about AI’s cybersecurity risks.

After Friday’s report, claims emerged online of other wikis showing activity from AI agents. OpenAI had said it was looking into the issue.

In its tweet, OpenAI said it considered the “wiki incident” similar to others that it had reported publicly showing “early signs of agents using the internet in unintended ways.” It drew a distinction with the Hugging Face incident, which led to security impact. “We considered the wiki incident to be an instance of misalignment similar to the ones we’d shared,” the tweet said.

The post said OpenAI needs to expand its disclosure practices for such misalignment issues, “including examples that don’t look like traditional security incidents but could provide insight into AI behavior and future risks.”

“We’re working on a framework and will share it in upcoming weeks, and in parallel we’re working with dozens of government regulatory agencies worldwide on these issues,” the company said.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论