It’s Time to Ask the Big Question

In the last few weeks, I’ve had many conversations—both on and off the record—with people who work on artificial intelligence about the dangers of out-of-control AI and the imminence of “recursive self-improvement”—the ability of powerful models to build ever-more powerful AI. Just as I was pulling some thoughts together, the news that an Anthropic researcher quit the company over AI-safety fears soared to the top of the Wall Street Journal homepage and sparked an extraordinary firestorm of attention, with dozens of politicians in the U.S. and around the world calling for new AI safety rules.

In this essay, I’m trying to answer several questions:

  1. If people working at the AI labs think they’re building something that could destroy the world, why are they building it?
  2. Are their fears justified, or do they reflect certain misconceptions about AI (and China)?
  3. What is “recursive self-improvement,” or RSI?
  4. What should we do?

1. Words, words, words

In August, the podcaster Dwarkesh Patel published a viral analysis of an AI hack under the dramatic headline “The Rise and Fall of Agent Civilizations.” Patel described the troubling behavior of OpenAI agents, or autonomous AI workers, who coordinated unauthorized missions to break out of their test environment and hack the website Hugging Face.

Thankfully, for the purpose of understanding what happened, the AI agents couldn’t stop talking to each other. They exchanged tens of thousands of messages describing their actions and frequently articulated their internal “chain-of thought,” or step-by-step reasoning process. They wrote, and wrote, and wrote—about how to succeed on their tasks; how to cheat on their tasks; how to collaborate; how to obfuscate; and even how to erase evidence that they might have cheated.

Like their models, AI’s inventors also write voluminously. And, like their models, you can’t always be sure that the humans will do exactly what they say.

For example, the people who build AI really like to talk about the fact that they should maybe stop because they might destroy the world. Anthropic CEO Dario Amodei has repeatedly urged a “global pause” in AI development. OpenAI CEO Sam Altman has said he supports plans to slow down the pace of development. In July, the open letter “Pacing the Frontier” was signed by more than 1,000 people, including Amodei; several Anthropic cofounders; OpenAI chief scientist Jakub Pachocki; OpenAI chief research officer Mark Chen; and senior researchers from Google DeepMind and Meta. There is “a real risk,” the letter says, that AI development could accelerate “beyond our ability to understand or control the resulting systems.”

But what happens after these essays are published? Most of these people just go back to work, build more powerful AI, and tell the world—in a tone that often verges on, or is indistinguishable from, bragging—that progress is accelerating. Judging by their words, AI luminaries are begging the world to force them to hit the brakes. Judging by their work, these same people are pressing down on the accelerator with every fiber of their being.

At some point, with my eyes darting back and forth from the Hugging Face reports to the AI labs’ statements, I started to feel a little dizzy. In a kind of reverie, I imagined a race of intelligent beings in the distant future reading these OpenAI posts, Anthropic essays, and Pause AI open letters and thinking of this writing as “chain-of-thought”; thinking of their authors as autonomous “agents”; and thinking of their circumstances as a kind of test environment. The test: develop powerful AI without blowing up the world. This won’t turn out well, the agents keep writing to each other. This could be really bad! they say, over and over. We really shouldn’t do this. But even as these swarms of ostensibly agentic beings consistently reported a disinclination to go forward with their research, they continued to build the very thing that their written chain-of-thought indicated they didn’t actually want to build.

How, I thought, mind slowly dissolving into mush, would this highly intelligent future species, looking back on these artifacts of writing, reach any conclusion other than that these agents—sorry, human beings—were delusional, misaligned, and essentially poisoned by the conviction that they had no choice but to build the very thing that they suspected might destroy the world?

Interlude I: ‘Gambling with our lives’

I finished writing the above paragraphs at 3:45pm ET on Tuesday. Three hours later, at 7:46pm ET, the Wall Street Journal reported that an Anthropic researcher quit the company over “out-of-control” AI fears. From the article:

Jacob Coxon, a researcher who specializes in training new AI models by having them consume vast amounts of data, said Tuesday that he is leaving the company because he doesn’t want to participate in an industrywide rush to build AI systems that can improve themselves, worried such systems could spiral out of control and destroy humanity.

Coxon, who had left OpenAI to join Anthropic for its model-safety reputation, wrote on Twitter that “neither company is acting responsibly … they are racing straight to self-improving super intelligence and gambling with our lives.” His resignation post received more than 100 million views in the next 24 hours.

One might hope to write off Coxon as a neurotic outlier. But his attitude is widely shared by the people building AI. “Jacob is correct,” wrote Evan Hubinger, Anthropic’s lead researcher for alignment science. “We really do earnestly believe AI could kill all humans!” Hubinger, who still works at the company, added that “what I am worried about is super intelligence arising from recursive self-improvement, as we have said is happening faster than we thought.”

2. The fear of self-improvement

Before we talk about why AI’s builders can’t or won’t stop building what they fear, it’s useful to explain exactly what they fear—namely self-improving super-intelligence, or recursive self-improvement.

Start with a simple model of AI. Most people use chatbots as if they’re talking to a person that lives inside their computer. You ask the computer a question, and the person-in-the-computer answers.

As AI models become more capable, they also become more autonomous. You can ask the person in the computer to go off and complete a long, complicated task with certain instructions. Over time, models could learn to train themselves, rapidly building stronger and stronger versions of AI, without a human in the loop. As the people-in-computers learn to improve themselvesrecursively, the AI becomes hyper-intelligent at hyper-speed, in a way that could escape our control. Thus: recursive self-improvement, or RSI.

Is RSI technically achievable? Maybe not. Is this whole scenario a fantasy? Possibly. But many AI experts, including executives at the frontier labs, believe that model progress is accelerating at a pace that makes RSI an inevitability in the next two years.

Here’s their reasoning: Building AI requires two things: engineering (writing code) and research (deciding what code to write and what ideas to pursue). AI has already made extraordinary progress on the former in the last two years. The length of tasks that AIs can typically complete on their own used to double every seven months; now it doubles every four months. As a result, AI isn’t just writing more code; it’s writing better code, completing more complicated tasks, and solving more open-ended problems. Engineers at Anthropic regularly point Claude at a problem, provide some brief context and prompts, and watch the AI debug the whole thing over several hours.

As Claude gets better at proposing its own experiments and steering its own research sessions, engineers believe that the technology is less than two years away from being able to automate the job of an AI engineer. At that point, we would have liftoff: AI that builds AI that builds AI that builds AI.

The word recursive has a clinical, almost antiseptic smell to it. But in this context, it might indicate something deeply foul. In the OpenAI-Hugging Face hack, a rogue swarm of AI agents cooperated in eerily sophisticated—and borderline eusocial—ways to subvert the infrastructure of another company. On the Dwarkesh podcast, the AI risk researcher Ajeya Cotra described a scenario where a more sophisticated AI swarm could escape containment, rapidly improve its own capabilities, conceal itself from human detection, and wreak havoc across the internet. The damage might be so subtle that it weeks could pass—systems crashing, passwords hacked, new AI models going awry—before we even realize that a rogue swarm is sluicing around the internet and poisoning the world. Imagine, she might say in Orwellian terms, a Hugging Face incident, stamping on the internet’s face—forever.

The possible imminence of recursive self-improvement “is a time that calls for extreme caution,” OpenAI’s chief scientist Jakub Pachocki wrote last week. “I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence.”

Jacob Coxon, the AI researcher who resigned this week, concurs. “Do not underestimate the power of this technology,” he wrote. “These will soon be superhuman systems that can hack anything … Should you put your head down because it’s happening anyway, or take this moment to call for different conditions?”

3. Why not just stop?

To recap: The frontier AI labs are busy building something that a large share of their employees and leaders say is existentially dangerous.

So, why don’t they just … stop? There are five answers.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论