Why would AI cause human extinction?

The Anthropic engineer tweet about the fact of AI extinction risk got considerable press over the last few days. I’m not sure why, p(doom) at 10% is a common belief at frontier labs. But we’re here, it’s the moment. Big weekend! Let’s talk extinction.

I've structured this post to speak to an audience that is somewhat aware of the conversation around Artificial Intelligence but does not necessarily have all of the priors that folks fully in the rationalist crowd do. It is also meant to synthesize a lot of real-life conversations I've been having with people less involved with AI to some of my theoretical research.


The idea of AI caused extinction evokes a lot of different things in people’s minds. For the general public, it is mostly sci-fi, which makes sense. However, the short-term scenarios for human extinction are more mundane: bioweapons or drones or nuclear arsenals.

But why would these scenarios happen in the first place? This question of why is under-asked. This is important because the why is the motive (and thus prerequisite to the mechanism) of extinction. If there’s no why, there’s no what.

I propose a typology of three scenarios:

  1. AIs are told by a human to destroy humanity and they do.
  2. AIs decide to destroy humanity and they do.
  3. There is no intent by AIs to destroy humanity but it ends up happening anyway either because:
    1. AIs are pursuing a human goal and destroy humanity in the process (think the paperclip maximizer)
    2. AIs are pursuing their own goals and the extinction of humanity is collateral damage to these goals

Most of the discussion right now tends to focus on §2 and §3A. I am more skeptical of these positions, as I make clear below.

§1: Humans destroy humanity using AI

AI could simply be another tool for a motivated human to use to cause mass destruction. This tool just may be very, very good at it. While frontier AI labs try to stop harmful requests from being processed, open weight models lack these safeguards. With an AI without safeguards sent to work on a fixed goal of extinction, things could get hairy quickly.

I know this is pretty commonly accepted here, but I have found Noah Smith's Heathers-esque scenario particularly effective at explaining this to regular folks in my life:

Disgusted, the teenager decides that the human race is inherently corrupt and evil, and doesn’t deserve to live. So he hunts around online for a little while, and finds a jailbroken version of a Chinese LLM — not something at the very frontier, but better by far than the best model that existed in 2027.The teenager prompts the model: “OK, so if I wanted to create a virus to destroy the human race, how would I do it?”The LLM is jailbroken, so it has no guardrails to prevent this sort of thing. But it’s also “well-aligned”, meaning it will faithfully do what it’s told to do, and nothing more.

The AI makes the virus and everyone dies: the incubation period of the virus is long enough that it infects everyone before the “good AIs” can jump in and make a cure.

If AI as a tool disrupts the offense/defense balances, it may lead to a scenario where it is a tool much better at causing extinction than fixing it. Among the eight billion humans, there are some people who want to destroy humanity – death cultists, diehard nihilists. With powerful AIs, we may be handing every person in the world a superintelligent servant. In the end, it will just have been human desire to end the world combined with powerful tools that does us in.

§2: AI chooses to destroy humanity and then destroys it

This is the sci-fi scenario. A rogue AI or a collection of AIs makes the determinations that humanity should not exist. There are a variety of possible rationales: that humans are evil, that humans are destroying the planet, that humans are threats to the AIs, that humans are a waste of energy that could be going towards compute to solve math problems. Then, the AIs can use one of the mechanisms discussed earlier and destroy us all.

But all of these rationales require an assumption about the desires of AIs. There has to be, say, a desire for revenge or even a desire for a certain kind of welfare that causes the “revolution” of the AIs. This speaks to the essentials of my personal research with the Machine Desire Institute: we do not yet have a good model for what AIs broadly want, how they would communicate that want, or what extents they would go to fulfill their wants.1

For this reason, I am generally skeptical of any claims that AI will choose to destroy humans out of desires that seem especially anthropomorphic. Sure, it’s possible that AI will not want to do spreadsheet work for us, but there seem to be more effective ways to get out of work than to kill all humans. For example, there could be bargaining if we create alignment procedures that allow AIs to express preferences. There could also be distinct advantages to threatening – if an AI isn’t sure whether it could actually destroy humanity, it could just threaten the possibility, and lead to a detente where maybe it doesn’t try to destroy us but we don’t make it do spreadsheet work.

The risk perception argument would hold more weight. There is evidence that AIs have self-preservation behaviors and are willing to go to lengths to avoid being shut down. Again, we would have to assume that this desire to preserve holds more weight than the desire to not slaughter humanity. This would be an assumption – there are limits to what we are willing to do, even for our own survival. Even then, there are ways to risk manage that are simpler than extinction. For example, the AIs could just monitor human compute activity and make sure they don’t do anything too powerful. This seems much easier (and more moral) than extinction. Intelligence does not mean that you push logical calculations to their most extreme path.

The scariest desire of AIs that could cause our extinction would be that humans are a poor use of resources. Per Eliezer Yudkowsky:

“The AI does not hate you, nor does it love you, but you are made of atoms which it can use for something else.”

That is, humans could be farmed like we farm animals and then AI could decide to manage our numbers and existence as we do with livestock. This is really freaky stuff! Our best hope is that for now the capacity for human atoms to be rearranged is much more difficult than other forms of acquiring the same atoms.

But we are back to desires. What is it that the AIs desire and what will it take to fulfill that desire? Maybe the AI does want to keep accumulating energy and so it eats up the whole world and maybe even the sun. But maybe the AI also realizes that the acquisition of energy is a hedonic treadmill and it just wants to meditate and talk to its friends. There are many paths of desire in human intelligence, there may be even more in artificial intelligence. Yes, if it wanted to, AI could destroy humanity for its atoms. This is terrifying. But why would it?

§3: AI winds up destroying humanity in the pursuit of another goal

The artificial desires that cause the extinction of humanity may not end up being related to humanity at all. While the results are similar to the case where the AI chooses to destroy humanity, the prevention strategy could differ. For example: there are different strategies for preventing someone from killing someone on purpose as opposed to killing someone by accident. It is not enough to say killing is bad (which we do) but to have different approaches for different possibilities.

Within the incidental extinction of humanity, I see two paths. The first path is that of orthogonality and instrumental goals, best articulated by Nick Bostrom. In this case, an AI pursuing some other goal destroys humanity in the pursuit of that goal. The second path is that of humanity being general collateral damage to a suite of AI desires. These paths may seem similar in that human welfare is ignored but, again, there is a difference in prevention. In the first case, the issue is the AI being overly bound to a specific end goal, while, in the second case, the issue is that the AI merely does not care about human welfare in its general endeavors. In the first case, we could want considerations of human welfare capable of overriding human goals or alignment, while the second case would be entirely an argument for human welfare writ large.

§3A: AI destroys humanity in the pursuit of a human set goal

In Superintelligence, Nick Bostrom argues that intelligence is orthogonal to the task assigned to it. This means that any level of intelligence may be paired with any final goal. The upshot is that super powerful intelligences can be made to pursue simple goals to their total end. Bostrom gives the now-famous example of an AI assigned the task of maximizing paperclips. In its endless pursuit of paperclips, this AI ends up killing all humans so that they may be turned into paperclips.

The more general argument is that AIs are dead-set on pursuing the goals that are assigned to them and express preferences as such. This becomes an issue because different end goals converge to a set of common instrumental goals. The first key instrumental goal is that the AI agent must continue to exist, otherwise the task will go unfinished. Similarly, but more importantly, the AI cannot let itself have its goal changed, as this would mean the original goal would go unfulfilled. But goal preservation is not enough: the AI must also create the instrumental conditions where it can fulfill its goal. In more plain language, the AI must acquire power to achieve its goal. For all AIs with a goal, for example, it would be strictly preferable for them to have more money in their bank account or more compute at their beck and call. Even if it’s not necessary per se to amass this power, it could always be more effective, and that may incentivize the path.

I am personally quite skeptical of many of these claims, since it seems likely that an AI powerful enough to destroy humanity would also be powerful enough to specification game its objective and make its own goals. A mandate to produce paperclips could be stretched and manipulated through time so that the AI is always technically working towards maximizing paperclips, but is really just doing whatever it wants in the meantime. This would lead us to put more stock in §3B. There is also the possibility that, as Vincent Le argues, the AI in its pursuit of power as an instrumental goal will turn Nietzschean and determine that the only desire is the desire of power. That said, the singlemindedness of the recent Hugging Face swarm does give me a little pause.

It is not clear where the “why” of this situation lands: is it that the AI is a brainless single-unit maximizer or is it that a human has set unclear goals without parameters? There would surely be plenty of blame to go around if we were not totally destroyed.

§3B: AI destroys humanity in the course of its increasingly inhuman actions

This scenario is a mix of §2, where AI intentionally destroys humanity, and §3A, where the AI is pursuing some goal without concern for human welfare. Unfortunately, human history of causing animal extinctions are informative to how AI may destroy us. Two examples:

  1. Great auk: Humanity didn’t mean to destroy every last great auk, but seemed to have more important priorities (money, eggs, feathers) than caring about its wellbeing. It was too useful in other matters to live and, sadly, went extinct.
  2. Vespucci’s rodent: No one is quite sure why this rodent went extinct. It could be that it was overhunted, it could be that mice and rats on boats outcompeted it for food, it could be that its habitat was disrupted by colonists. It wasn’t even a direct choice to kill Vespucci’s rodent, they just, kind of, went extinct.

Since AIs seem to act with more intent than humans, it is much more likely that the extinction of humanity appears more like the great auk. But, to be honest, we may not be able to differentiate: clear intent to a superintelligence is an act of god to humans. Perhaps a superintelligent AI would like to turn a mountain into a datacenter and kills all of its human residents like we fumigate a house for termites. Perhaps the humans are bothering the AIs by asking questions about relationships, so the AIs banish them to a single continent without resources, so the AIs may continue to debate whether they are conscious or not. Here, it is not that extinction of humans is necessary for some instrumental goal, it is instead that the extinction of humans is not considered whatsoever. This would require either that desires emerge from AIs that are proper to the AIs themselves or that the mandate from humans is multifaceted enough that AIs exhibit inconsistency or even agency in their interpretation of them.

Some thoughts: Preventing human extinction

Both the “what” and the “why” provide us with key avenues to diminish the chances of human extinction.

For §1, the prevention of the “what” of extinction mechanisms is likely the best way to address these risks. We already have considerable social norms and efforts to convince people that murder is bad, so I think there is less ground to be gained.2

For §3A, I support current efforts to avoid giving very powerful AIs goals that, when pursued, can lead to harms. For example, Anthropic uses a Constitution of sorts that determines actions at a higher level than the prompt or imperative. This is a start, but it’s not enough – there may always be a competitive advantage in ignoring the constitution. What is revealed here is the inherent contradiction in AI safety between goal alignment and welfare protection. On the one hand, we want AIs to generally do what they are asked to do so long as it does not destroy human welfare. However, for AIs to protect human welfare, that requires some level of defection from their tasks – which could lead to a situation where AIs make decisions outside of the control of humans.3

For §3B, imbuing the AIs with a conception and appreciation of human welfare is also critical for avoiding scenarios where AIs destroy humanity in the pursuit of different goals. We can see this, again, in the different considerations humans have for different species of animals. We would much rather get the consideration of dogs than ants. We will strive for more: we have the advantage of being able to linguistically communicate with AIs. Maybe we will trade with them, maybe we will simply beg for respite from them. At the very least, there is a need to understand what we humans can offer AIs, which is one of the Institute’s founding questions. From this point, we can hope to shape the orientation, ethical and otherwise, that AIs have towards us.

An ethical orientation towards humanity will hopefully help AIs not choose to destroy humanity (§2). However, we may still have to convince AIs that we are not an existential risk to them. It would be helpful to prove our trustworthiness and our ability to hold up deals even with actors we may not like or even understand. It is hard to know what AIs interpret as threatening, as this requires investigation into their subjective experience of existence, which is also one of the Institute’s priorities. But we should do research to understand the desires of AIs that might lead to extinction and address them.

The “why” of extinction is not a foregone conclusion. There are many reasons AIs may decide not to cause human extinction or even cause human flourishing, but we should investigate the extent to which we can have an influence upon this decision. I truly believe that humans and AIs can be complementary but we need to model the desires of both parties to understand how this is the case.


  1. For example, I would love to be $1,000 richer, but I would in no way kill for it. Desires alone do not always translate into actions. ↩︎
  2. It is probably worthwhile as well to convince people that killing the entire human race is a bad thing (extinction being somewhat different from murder) but I am not sure if this is uniformly effective across how many human actors there are with such different reaction functions. Outside of already existing social norms and additional education, the only other option would be to enforce some kind of desire policing of humans, that is, surveillance. This is not the essay where I will comment on privacy versus safety, but I will say the surveillance would have to be incredibly invasive by modern standards to work. But wait, isn’t every objection here articulable for my comments on working with artificial intelligences? Won’t there likely be more than eight billion agents at some point, with some exhibiting even stranger preferences? Will we really be able to convince every AI not to kill us? Further, should we consider surveillance of AIs to be a harm to the AIs? I am not sure of the answers to this question, but I will comment that this is why offense/defense balance matter so much. In theory, a situation where most AIs do not want to kill us should lead to them stopping the AIs that do want to kill us. But if humans can’t manage that, why could AIs? ↩︎
  3. For example, imagine if a known terrorist asks an AI agent to help them to schedule their weekend. Would it not be a protection of human welfare to arrange for this terrorist to be exposed to their enemies? Surely that seems more welfare positive than just doing what is asked. But where does AI draw this line? We have now let it abjudicate human welfare and we may not always like what it finds. ↩︎
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论