What pacing the frontier means for China
Dario's recent post on pacing the frontier mentioned how slowing down China more would help with international coordination. It seemed counter-intuitive, surely sabotaging their efforts would create badwill, but it may give the US more leverage to negotiate. Would China be hyper-rational about it, or would hurt feelings make coordination tougher?
Even today, looking at smart people responding so differently to AI risk seems very counter-intuitive. Something related to incentives and wishful thinking.
Here are some quick explorations with some numbers.
The objective function
Let us define reward R as the objective function that each person tries to maximize. Let's assume we maximize the EV of R.
I can think of 2 reasons R is different for different people :
- They want to work towards different things.
- They want to work towards the same thing but reason about likelihoods differently. One of them is likely more correct and the other is more irrational in their estimates.
An example of 1 could be someone mission driven (net good including others > wealth) vs someone chasing money or status.
An example of 2 could be someone engaging in wishful thinking or denial or not being well informed enough. For example, the likelihood estimates in ai-2040.com are probably much more well informed than by someone with limited context.
Can someone who cares about the world be corrupted by incentives?
Dario and Sam may have a very similar final goal, something along the lines of "Humanity that thrives in an age of abundance in a democratic world."
However incentives shape how they calculate their likelihood estimates. For example :
- If OpenAI has it's IPO coming up, and not enough people inside OpenAI voice concerns about AI safety, the likelihood in Sam's head about AI going badly would be lower. A super-intelligent, hyper-rational human would not be prey to this but we're not super-intelligent and hyper-rational.
Similarly a lot of VCs who have incentives for an open AI ecosystem, want to see startups they are investing in thrive. They may be engaging in wishful thinking that reduces their likelihood estimates of AI going badly. Or they may feel averse to reading the relevant AI safety literature.
However organizations and nations may be much more rational than individuals, owing to the diversity of perspectives and experts involved.
The US China situation
If the US tries to slow China down, will badwill reduce the chance that we coordinate?
There are examples where emotional reasoning led to bad decisions by states.
- Japan, 1941. After the US oil embargo, Japan's planners estimated it would love a long war with the US. It still attacked Pearl Harbor, because accepting the embargo's terms read as national humiliation.
- Austria-Hungary, 1914. The assassination of Franz Ferdinand was evidence about a Serbian nationalist network. But it turned into a war to punish Serbia.
- The US after 9/11, Iraq 2003. The 9/11 attack was evidence about al-Qaeda. It soon became the lens every threat was seen through, and it led to a lot of anti-Iraq sentiment.
But it seems in this case, that having some chip export controls will not be seen by China as humiliating or deserving of military retaliation. So let's model China as a rational player.
The incentive matrices
Each side has two moves. Trust means following the agreement. Hedge means secretly go faster.
Hedging can mean going faster without racing at full speed to catch up. However if one side races, the other will too.
Dario's post has two kinds of agreements:
- Level 3 : Some kind of pacing agreement to reduce speed of progress.
- Level 4 : A full pacing, or a pause. Where we pace conservatively.
Reward Equations
Let's model normalized rewards as :
R_US = 𝟙[humans in control] × 𝟙[liberal order survives] × (0.7 × US power + 0.3 × abundance for it's people)
R_China = 𝟙[humans in control] × 𝟙[regime survives] × (0.6 × China power + 0.4 × abundance for it's people)
With 𝟙[humans in control] meaning 1 if humans are in control and 0 if not.
Fixed assumptions:
- Abundance for people is 0.9 in every world where humans survive, for both countries.
- Inspired from estimates in AI 2040, estimated P(humans stay in control) :
- 85% if both pause.
- 75% if both pace under a conservative speed limit.
- 60% if one side secretly goes faster.
- 50% if both do, or if there is no deal.
- 30% if anyone races.
- Power is 60/40 today and stays there under an agreement. Cheating under a speed limit gets you 75/25. Cheating under a pause gets you 1/0.
- If someone wins outright, the loser's system survives only if the winner holds back. China thinks the US would hold back 40% of the time. The US thinks China would 30% of the time. Otherwise each system survives 95%.
- A cheater gets caught 10% of the time without verification, 90% with verification measures in the deal.
Here are the reward EVs in different situations:
Outcome | Cell | R_US | R_China |
|---|---|---|---|
Both pause, honest (85%, 60/40) | TT (pause) | 0.85 × 0.95 × 0.69 = 0.557 | 0.85 × 0.95 × 0.60 = 0.485 |
Both pace, honest (75%, 60/40) | TT (speed limit) | 0.75 × 0.95 × 0.69 = 0.492 | 0.75 × 0.95 × 0.60 = 0.427 |
Both secretly faster (50%, 60/40) | HH | 0.50 × 0.95 × 0.69 = 0.328 | 0.50 × 0.95 × 0.60 = 0.285 |
US ahead, speed limit (60%, 75/25) | US cheats, not caught | 0.60 × 0.95 × 0.795 = 0.453 | 0.60 × 0.95 × 0.51 = 0.291 |
China ahead, speed limit (60%, 25/75) | China cheats, not caught | 0.60 × 0.95 × 0.445 = 0.254 | 0.60 × 0.95 × 0.81 = 0.462 |
US wins, pause (60%, 1/0) | US cheats, not caught | 0.60 × 1 × 0.97 = 0.582 | 0.60 × 0.4 × 0.36 = 0.086 |
China wins, pause (60%, 0/1) | China cheats, not caught | 0.60 × 0.3 × 0.27 = 0.049 | 0.60 × 1 × 0.96 = 0.576 |
No deal, US export controls (50%, 65/35) | Cheater caught (speed limit) | 0.50 × 0.95 × 0.725 = 0.344 | 0.50 × 0.95 × 0.57 = 0.271 |
Everyone races (30%) | Cheater caught (pause) | 0.30 × [0.6×0.97 + 0.4×0.3×0.27] = 0.184 | 0.30 × [0.6×0.4×0.36 + 0.4×0.96] = 0.141 |
TH and HT = catch rate × caught row + (1 − catch rate) × not-caught row.
Calculating Equilibrium
Let's treat this as a mixed-equilibrium problem to solve that has a formula. Let's call trust T and hedge H.
P(US hedges) = (HT − TT) / [(HT − TT) − (HH − TH)] using China's payoffs
P(China hedges) = (HT − TT) / [(HT − TT) − (HH − TH)] using US payoffs
If either comes out below 0 or above 1, that side has a dominant move and the answer is 0% or 100%.
Now let's run some of the cases.
Case 1: Pause, no verification
Cheaters get caught 10% of the time. If caught, the other side sprints.
R_US: TT 0.557, TH 0.062, HT 0.542, HH 0.328 R_China: TT 0.485, TH 0.092, HT 0.533, HH 0.285
China hedges no matter what. Because they can pace in secret, keep their AIs aligned, not get caught.
China: Trust | China: Hedge | |
|---|---|---|
US: Trust | 0% | 0% |
US: Hedge | 0% | 100% |
Survival 50%.
Case 2: Pause, verified
Same pause, with compute accounting and inspectors. Catch rate 90%.
R_US: TT 0.557, TH 0.171, HT 0.224, HH 0.328
R_China: TT 0.485, TH 0.136, HT 0.185, HH 0.285
China: Trust | China: Hedge | |
|---|---|---|
US: Trust | 100% | 0% |
US: Hedge | 0% | 0% |
Survival 85%.
Case 3: Speed limit, no verification
Both keep building under a cap on RSI speed. Catch rate 10%. If caught, the deal dies and both pace faster with no agreement.
R_US: TT 0.492, TH 0.263, HT 0.442, HH 0.328
R_China: TT 0.427, TH 0.289, HT 0.443, HH 0.285
China: Trust | China: Hedge | |
|---|---|---|
US: Trust | 11% | 9% |
US: Hedge | 46% | 34% |
Survival: 0.11×75 + 0.09×60 + 0.46×60 + 0.34×50 = 58%. Messy but better than Case 1.
Case 4: Speed limit, verified
Same speed limit. Catch rate 90%.
R_US: TT 0.492, TH 0.335, HT 0.355, HH 0.328
R_China: TT 0.427, TH 0.273, HT 0.290, HH 0.285
China: Trust | China: Hedge | |
|---|---|---|
US: Trust | 100% | 0% |
US: Hedge | 0% | 0% |
Both comply. Survival 75%.
Case 5: No deal, the US attacks China, China's feelings are hurt
Let's assume the US paces the frontier on it's own. Then it strikes Chinese AI infrastructure to widen its lead. Some chinese people die.
What China's incentives say it should do
China's move | R_China |
|---|---|
Keep pacing, no deal | 0.271 |
Get a verified speed limit | 0.427 |
Race | 0.141 |
It should ask for a deal even more now, since the US has a bigger lead and without a deal the US would win harder.
But China is hurt. It will not race, that is fairly suicidal. But there won't be a deal and it would likely pace faster than it would have without an attack.
Case | R_US | R_China | Survival |
|---|---|---|---|
1. Pause, no verification | 0.328 | 0.285 | 50% |
2. Pause, verified | 0.557 | 0.485 | 85% |
3. Speed limit, no verification | ~0.38 | ~0.35 | 58% |
4. Speed limit, verified | 0.492 | 0.427 | 75% |
5. No deal, US strike, China hurt | 0.184 | 0.141 | 40% (no deal case but faster pacing) |
What does this suggest?
- We may not need to worry about badwill with China unless we attack them. The incentive structure is too strong.
- With less verification, they will likely pace enough in secret to catch up with the US, again as long as there is not high risk of takeover.
- A bigger lead means they are less likely to break a treaty to catch up, because catching up is tougher.
- They will also likely pace the frontier as long as they realise takeover risk is real, which they likely do.
The US should probably :
- Be very vocal about AI risk and the actual estimates so China can anchor on them.
- Try to get a bigger lead without sabotaging China harmfully.
This is all, only insofar as alignment can be done successfully with the current pacing that's being proposed. If it cannot then a pause treaty becomes the only way out and much more likely.