Plan A, by AI-2040

The folks who wrote AI 2027 have written a more optimistic narrative, which focuses more on hopes for good policies than on predictions about what policies we’ll get.
Plan A’s narrative seems halfway between a science fiction story and a proposed treaty. Like most science fiction, I expect it to err in the direction of describing the world as more human-understandable and relatable than what we’ll actually get.
The broad outlines come close to the scenario that I analyzed in Financial Costs of an AI Pause?, which is what I predict that fairly competent governments would do.
AI-2040 adds much more detail than I was able to provide, some of it surprising. The devil is in the details.
I largely endorse their advice. The rest of this post will focus on many small doubts about their advice and their predictions about what that advice would produce.
Keep in mind that this is just a plan. Expect plans to change in response to contact with reality. Rates of AI capability growth ought to change in response to better evidence about the difficulty of alignment.
Please read the Insider Perspective section. It’s a little more technical, but it answers several nontechnical questions that were covered inadequately in the main story.
What Success Looks Like
The long-term vision of Plan A is to hand control over to AIs that we’re pretty sure have our interests at heart. The authors guess that will happen in 2040.
Plan A’s vibes suggest the handover will go well, but the authors seem careful to avoid saying that such a decision would be wise.
Most of the plan describes global agreements to slow AI progress, by enough that safety research will have time to be more thorough. That includes a near total pause in increased AI cognitive abilities around 2035, when AI matches the abilities of top human experts at almost all tasks.
The authors make some assumptions about AI that are mildly controversial. I expect those assumptions will turn out to be fairly reasonable.
They describe a world with lots of semi-equal AIs, with diverse alignment targets. They’re not necessarily claiming that we’d get this diversity in the absence of regulation, but the authors seem rather confident that Plan A will produce a multipolar scenario.
I’m something like 70% confident that they’re correct. It seems important to flag this as controversial. A multipolar world ought to be achievable, but may require better plans than Plan A has articulated.
They forecast reliable lie detection for both AIs and humans. That enables alignment of AIs, politicians, CEOs, etc.
2038: AI Alignment Is Now a Science Want your new AI to be honest? There’s standard protocol for training true honesty … Different AIs have different alignment targets programmed in.173 The reason this says “programmed” instead of “trained” is that the rapid advances in alignment have resulted in alignment techniques that can directly program in an AI’s goals by directly modifying the AI’s code/weights, unlike the alignment techniques of 2026.
In spite of massive restrictions on economic growth, the US output growth accelerates to a bit over 80%/year in 2033, and stabilizes at that rate through 2039. Presumably it accelerates beyond this growth rate after restrictions are lifted in 2040. These forecasts seem high to me, but still more plausible than a majority of the forecasts that I see for the effects of AI. I’m guessing more like 30 to 60% per year growth given what I expect to be the default restrictions, which will likely be weaker than what the authors want.
Here’s a graph that they provide to compare this growth with historical trends. Pay close attention to the unusual x-axis:
Stability?
Would governments agree to this deal and enforce it?
How strong are the incentives to defect? Is this a winner takes all race? Or does finishing in third place doom a group to merely being billionaires, while the winners are trillionaires?
The arguments against a deal don’t seem obviously stronger than arguments in the late 1940s against a deal to limit nuclear weapons. We’re further from a cold war than the world was in the late 1940s, and the Baruch Plan wasn’t hopelessly far from working. Still, the time it took for an actual nuclear deal is grounds for concern.
Politicians aren’t taking this very seriously yet, but opinions are shifting enough that current opinions don’t tell us much about next year.
We seem approximately on track for increasing fire alarms that are roughly what it will take to scare governments into an agreement that resembles Plan A, but the situation is sufficiently unusual that my predictions are pretty low-confidence.
Freeing Thucydides offers hope that the deal will become more stable over time:
The Thucydides trap is a particular manifestation of commitment failures: a rising state cannot bind its future self, so the declining side cannot trust its promises, and may prefer to fight while it still can. … Powerful AI systems could provide the enforcement mechanisms.
AI enforcement abilities will be pretty weak when Plan A needs the deal to start, but those abilities will grow significantly.
The deal becomes more stable once human lie detection works well:
International agreements are now more stable than ever, because it’s so hard to cheat. Anyone deciding to cheat needs an excuse for not being willing to prove that they aren’t cheating.
One place where the authors appear overconfident is this plan associated with 2035:
Making deals with misaligned AIs: a third line of defense
Deals with AIs are worth attempting, but why do we expect AIs of 2035 to keep their word? Do these AIs have enough continuity for promises to be meaningful?
AI-2040 has companies being pressured in 2033 to train “truthseeking AIs”, but how effective is that training? Are those the AIs we’ll want to make deals with?
AI-2040 has AI lie detectors becoming reliable in 2037, so there’s some hope of this being a temporary problem.
Tom Davidson presents concerns about allowing significant compute growth:
Riskier intelligence explosion. If the deal breaks down and there’s a race to superintelligence, it will be much faster and more dangerous than if we’d never done Plan A. … if the default trajectory is that there’s no software-driven intelligence explosion (because of compute bottlenecks) and extinction risk is 10%, then dry tinder can make things much worse. It creates a fast intelligence explosion (where there otherwise wouldn’t have been one) and dramatically raises AI takeover risk.
This assumes that the speed of the explosion is a major cause of risk. That’s not at all clear. I expect the risks to be influenced by our knowledge of how to align AIs at the start of the explosion. Since I expect incremental progress in that knowledge, that effect works in the opposite direction from the increasing risks of a faster explosion.
The explosion doesn’t automatically go to the maximum feasible speed. Both AIs and humans have some influence over the speed, and the incentives are complicated.
In sum, I see plenty of uncertainty as to how risky it is to allow compute growth.
Robot Cap-and-Trade
To solve these problems, the Consortium countries agree to restrict AI-enabled industry to special economic zones (SEZs) subject to similar transparency and monitoring schemes as the datacenters, and to cap their total robot and compute production at ‘only’ 4x annual growth. … In 2032, the US has a cap of 80M robots and 5 billion H100-equivalent GPUs. The market is so desperate for more robots and compute that permits become the expensive binding constraint, costing on the order of $200k per robot permit
I expect the robot part of this to significantly slow economic growth. $200k per permit seems lower than what I’d expect given Plan A’s assumptions.
This will be the hardest part of the deal to negotiate. I don’t see a Schelling point for balancing the interests of China and the US. A per capita limit would reduce China’s lead over the US. Whereas the US would object to basing the caps on whatever numbers a…