The AI race is already multipolar

TL;DR: Collaborating to reduce catastrophic risks seems very possible even for policymakers with very different goals.

The race for general superintelligence is often described as a bipolar race between two rival hegemons, the US and China. A lot has been written about how, even within this framework, it is not inevitable that both countries race ahead with minimal restraints.

But more fundamentally, and in the tradition of Box, I think this two-outcome model is so broken and unhelpful that it’s worth critically examining and hopefully replacing with a more useful one.

I'd been sure that something like this would have been written down in this forum before, but I haven't found it after some searching. So here’s my attempt at a better model: It’s a “race” between (at least) eight outcomes, not two.

Why eight? I’m trying to cleave the “outcomes” of the race such that:

  1. They’re (as much as possible) disjoint outcomes.
  2. They describe lasting endpoints that are stable after we’ve arrived there.

Here are the outcomes I find most plausible.

The Outcomes

Outcome 1: The United States builds superintelligence aligned with its own values.


Outcome 2: China builds aligned superintelligence aligned with its own values.


Outcome 3: A company, some other form of private institution, or government different from these two builds superintelligence aligned with its own values.


Outcome 4: A large set of multinational or global actors collectively develops superintelligence aligned with the values endorsed by the large majority of actors at the table.

Outcome 5: There’s a large set of powerful / superhuman AI actors with competing priorities, in ways that nonetheless are collectively broadly aligned with human values.

Outcome 6: A powerful enough AI or sets of AIs is built by some actor such that human judgment and control is lost, as described in the gradual disempowerment literature.

Outcome 7: “Fast takeoff”-type loss of control, leading to catastrophic outcomes.

Outcome 8: There is lasting and roughly stable consensus that building superintelligence would be harmful, actors who disagree have sufficient incentives to avoid pursuit of superintelligence, and it doesn't end up being built. We can also add “Superintelligence is conceptually impossible” into this category. Either way, it doesn’t happen.

In fact, there’s a ninth outcome: none of the above; none of the taxonomies described above meaningfully describe the outcome that ends up happening. This is a very live consideration, and becomes more live as time progresses (more on that later), but given that this is the “unknown unknowns” bucket, I can’t really go into more depth here.

Within the Outcomes

We can zoom into any of these individual slices, and within them they also hold different possible worlds with different levels of goodness.

For example, there might be an outcome 1a world where the US and global distribution of political and/or economic power becomes more unequal, or an outcome 1b world where political and/or economic power diffuses among more actors.

Each of these possible worlds would have different benefits, risks and key challenges. Particularly, even in “aligned” timelines that are short-term, it is contested that even in worlds where we have solved the technical alignment problem, that by default we get outcomes that make most people better off, and seems very unlikely that we get outcomes where a large percentage of the world’s population has had (or chosen not to have) meaningful input on the process.

Another important thing to notice within this framework is that for most actors, at least five of these outcomes are very bad. Particularly, it should be possible, even for actors with completely disjoint sets of acceptable outcomes (say, an American policymaker that thinks only 1 and 4 are acceptable, and a Chinese policymaker who thinks only 2, 3 and 5 are acceptable), to work together to reduce the likelihood of outcomes they both find completely unacceptable.

Outcome Likelihood and Value Depends Substantially on Time of Arrival

A final thing within this framework I find useful is to think of these slices as conditional distributions with respect to time.

This is useful if you expect alignment quality and success to be very time-dependent, as I do. I don’t think anyone knows how to build superintelligence safely as of 2026. I think it’s highly likely that this remains true throughout the 2020s.

It's likely that even if we understood it conceptually, we’d spend a lot of time determining implementation challenges, and let's not fail to mention, determining the extent to which societies even want or do not want superintelligence.

The likelihood distributions within the slices also change over time, as well as the value of each future conditional on being within these slices, to the extent that it’s possible to define a sense of “objective value” on a future.

Obvious caveat that all of these credences are gestures at gestalts. They’re inevitably approximations. As I dig into this framework more I hope to develop a more granular, grounded personal sense for some of the relevant probabilities, and if people find this framework useful I hope more qualified modelers develop and share their own, too.

Example Slices

Concretely, here’s my top-line slice conditional on superintelligence in 2027:

Outcome 1 - US ASI: 4%

Outcome 2 - China ASI: 1.5%

Outcome 3 - Private / Non-US/China ASI: 5%

Outcome 4 - Large body ASI: ~0.00001%

Outcome 5 - Collectives of Aligned ASIs: ~0.01%

Outcome 6 - Gradual disempowerment: 55%

Outcome 7 - Misaligned fast takeoff: 25%

Outcome 8 - Moratorium / Ban : 0%

Outcome 9 - None of the Above: 9.5%

Here’s my top-line slice conditional on superintelligence in 2040:

Outcome 1 - US ASI: 8%

Outcome 2 - China ASI: 6%

Outcome 3: - Private / Non-US/China ASI: 2%

Outcome 4 - Large Deliberative Body ASI: 16%

Outcome 5 - Collectives of Aligned ASIs: 12%

Outcome 6 - Gradual Disempowerment: 26%

Outcome 7 - Misaligned Fast Takeoff: 5%

Outcome 8 - Moratorium / Ban: 0% (Remember, we’re conditioning on superintelligence. This branch is very much live but not reflected in this conditional distribution)

Outcome 9 - None of the Above: 25%

Here’s my 2nd-layer slice conditional on gradual disempowerment in 2030:

Outcome 6a - Total disempowerment: 70% (not necessarily immediately, I see this branch more like “competitive pressures ultimately move too quickly for humans to reject disempowerment)

Outcome 6b - Relative disempowerment: 20%

Outcome 6c - Other: 10%

Here’s my 2nd-layer slice conditional on gradual disempowerment in 2040:

Outcome 6a - Total disempowerment: 30%

Outcome 6b - Relative disempowerment: 45%

Outcome 6c - Other: 25%

Be cautious that some of these layers are very, very unlikely, so importing intuitions from the overall unconditional distribution onto a lower layer is very unwise; my intuition is that a world where we delay the development of superintelligence by many years compared to the default path is a much wiser world than it appears from our current 2026 vantage point, so we should try to import more optimistic intuitions when reasoning about these better outcomes.

Not to mention, these things are just going to be very hard to reason about, so a lot of humility is warranted, in general. Things will likely become clearer / easier to reason about as outcomes and possibilities either get closer or get ruled out.

If we group outcomes 1-5 as “alignment-successful” outcomes, and outcomes 6-7 as misaligned outcomes, and if you buy my credences directionally in terms of “misalignment less likely as time progresses”, then delaying superintelligence is less about shifting the relative likelihoods for a US vs China victory and more about shifting from broadly misaligned outcomes:

To less misaligned outcomes:

That doesn’t mean harder geopolitical discussions won’t have to happen later, but it means we can make progress now towards the policy goal of reducing the likelihood of catastrophic outcomes from attempted superintelligence development.

  1. Analogous in some sense to the concept of values lock-in.
  2. For example, if I had to guess, the third most likely country to develop superintelligence over the next 15 years is South Korea.
  3. They’re also internal models, not all-things-considered probabilities.
  4. Specifically, critically misaligned; it's likely that the models will not be perfectly or provably aligned; we're trying to capture the subset of cases that are so misaligned that they lead to catastrophic outcomes.
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论