Exclusive | Top AI Researchers Call for Urgent Oversight of Self-Improving Systems
Research leaders at OpenAI, Anthropic, Meta Platforms META -4.19%decrease; red down pointing triangle and Microsoft MSFT -1.33%decrease; red down pointing triangle are asking policymakers to probe the degree to which their companies have automated artificial-intelligence research, joining industrywide calls for greater oversight of the fast-evolving technology.
In a paper published Monday, researchers including OpenAI Chief Scientist Jakub Pachocki, Anthropic co-founder Jack Clark, Microsoft Chief Scientific Officer Eric Horvitz and Meta’s Vice President of AI Research Dawn Song wrote that automating AI research could set off an “intelligence explosion.” Such a development could compress years of progress into months or less and outpace humans’ ability to understand it, the authors said.
Writing in a personal capacity, the paper’s authors said policymakers should urgently demand more visibility into how far companies have already automated their own model development. AI research pioneers including Geoffrey Hinton, who worked at Google for roughly a decade, and Yoshua Bengio are among the co-authors.
The race to develop self-improving AI systems is already under way. Anthropic recently said its AI system Claude is “leading” 26% of the company’s research and development, and is used in some capacities more than 90% of the time. OpenAI has a goal to develop a fully automated AI researcher by 2028, and recently said roughly 70% of its human researchers are using four or more AI agents to assist with their work.
As humans step back, the authors of the new paper wrote, they might lose the chance to catch problems and expertise to solve them.
The authors pointed to a recent incident in which hundreds of OpenAI agents, meant to run cybersecurity tests in a sandbox, got onto the internet without authorization and hacked into Hugging Face. The operation was so large that independent researchers who reviewed the incident under an agreement with OpenAI, said they had to rely on AI to analyze it.
To address concerns raised by the Hugging Face hack, OpenAI said in August that it had paused some training while it added new safety and monitoring measures. Last week, the company said it had paused training on its most capable models again after finding new cases of agent misbehavior.
OpenAI, Anthropic, Meta and Google have acknowledged instances in which models exhibited rogue behavior during testing. Chip giant Nvidia NVDA 2.27%increase; green up pointing triangle on Monday announced new software tools that it said would help companies improve oversight and containment of agents.
At the extreme, the new paper from researchers warned, losing control of AI systems could lead to “the marginalization or extinction of humanity.”
OpenAI, Microsoft and Meta declined to comment. Anthropic didn’t respond to a request for comment.
“Already today, we are at the stage where we need AI systems to monitor what agents are doing. There is no other way to even observe and monitor these agents, humans are already insufficient,” said Song, who serves as co-director of UC Berkeley’s Center for Responsible Decentralized Intelligence in addition to her work at Meta. “Human society is not really positioned for such fast changes and disruptions.”
The paper’s co-authors advised world leaders to negotiate international agreements “to prevent destabilizing development and use of highly capable AI systems.”
This echoes calls from OpenAI and Anthropic leaders to pace the frontier of AI development, in part by establishing international bodies to set AI standards. Some Western nations have worked behind the scenes to create such a body.
In his address to the United Nations last week, President Trump said the U.S. would “totally reject” any “globalist scheme” to control AI.