Assessing the impact of safety work needs equilibrium analysis (now more than ever)
TLDR: This post explains two equilibria which regulate the level of AI safety: The first describes how much resources AI companies are willing to spend on AI safety work due to commercial incentives. The second one is about risk awareness and most notably affects government interventions for safety. Doing safety work similar to what AI companies do usually doesn't shift the equilibria much, whereas other work, like more ambitious safety approaches or policy advocacy, do shift them. Both equilibria have become far more important recently, after the Hugging Face (and similar) incidents.
The equilibrium of commercial safety interests
Consider this simplified model:
- AI companies have commercial incentives to invest in safety research: it improves their brand and prevents their AIs from causing harm that triggers lawsuits or regulation. Therefore they will fund safety work until the marginal commercial benefit of investing a dollar in safety equals the marginal commercial benefit of investing a dollar in AI capabilities.
- Thus, if you're at an AI company doing commercially-incentivized safety work, e.g. training models to not take harmful actions, the counterfactual impact (henceforth just "impact") of the safety work you produce is roughly zero because it would've been done anyway. (Though if you're someone who actually cares about safety it's likely you're doing a better job than just what would've been commercially incentivized). There's even a small negative externality if you're a high-powered talent and the company would have to pay more to find a similarly competent non-EA person to do your work equally well, because then you're freeing up some AI company capital which they likely mostly use to race faster.
- The same applies if, outside of an AI company, you're doing safety work that's commercially useful to AI companies. Some safety work is more specifically useful for avoiding xrisk and less useful for protecting commercial interests, and the model applies much less to such work.)
- This applies analogously for funding safety work.
FAQ
Q: Does that mean working in safety at AI companies is useless or slightly harmful?
A: No, although some work could certainly be harmful. I think it's important that safety work is done by people who actually care about existential safety rather than people goodharting on legible safety metrics, in particular to make sure AI companies don't just paper over problems in ways that produce deceptive AIs.
It's also plausible that some AI companies spend some money on existential safety that isn't just a side effect of commercial incentives.
I also think there are other fine reasons to work at AI companies, e.g. for earning to give or for trying to improve the sanity of people you work with. (There are also more reasons against, but the details are beside the point of this post.)
Q: So do you think people have been funding the wrong work?
A: Well yes, I think there have been a lot of bad decisions in how funding was distributed, but no, I think the failure mode from not taking into account this particular equilibrium dynamic hasn't played a very large role so far.
That might be partially because until recently, commercial incentives for work that is at all existentially relevant were quite low. I think now the commercial incentives are higher, and it's plausible to me that they rise even more. So it's useful to keep in mind that funding safety work that looks like it could maybe also come out of an AI lab in a couple months is much less useful.
Q: Can't I just use counterfactual reasoning instead of thinking about equilibrium dynamics for evaluating impact?
A: Yes, if you consider the counterfactual properly. I think counterfactuals are often considered too narrowly though, and I find the equilibrium frame generally useful for better understanding the world. Though I think both are important.
Now let's look at another equilibrium that I think is even more important.
The risk awareness equilibrium
Since the Hugging Face incident, companies see much stronger commercial incentives to invest in safety and security. In other words, we shifted from an equilibrium where commercial safety interests were low to an equilibrium where they are significantly higher.
This change was largely caused by a change in the risk awareness of companies, which is a variable which itself lies in a larger equilibrium, right alongside the risk awareness of governments and of the public:
- Warning signs cause risk awareness which causes more safety effort which causes fewer warning signs which causes lower risk awareness which causes less safety effort. Safety effort is regulated like a thermostat around the equilibrium where the safety effort seems adequate to the warning signs people see (which doesn't at all imply that the actual safety is adequate, e.g. consider AIs scheming to seem nice).
- ("More safety effort" is to be interpreted broadly. Governments mostly don't increase safety effort directly but pass regulations which have that effect. A treaty is a special case here that strongly increases safety effort.)
- Admittedly, this equilibration isn't very smooth, i.e. the thermostat-like regulation can be delayed and jerky. Large warning signs like the Hugging Face incident can appear suddenly if early problems stay undetected or there's a sudden transition in the behavior of an AI swarm. And government action may undershoot or overshoot the equilibrium. So we shouldn't assume the current state is actually at the equilibrium, but in expectation things are still moving towards the equilibrium.
So unless the AI safety charity community is rich enough to put more money into safety effort than would happen at the equilibrium point, funding safety work that reduces warning signs may be mostly useless according to this model, because it just substitutes for effort that would've otherwise been spent anyway due to more visible warning signs.
Of course, there are kinds of safety efforts that don't significantly affect the probability or visibility of warning signs, like ambitious alignment moonshots or treaty verification mechanism research. This consideration isn't an argument against funding such research.
Interestingly, what still seems useful to fund according to this equilibrium seems pretty similar to what seems useful according to the commercial safety interests equilibrium we considered before, perhaps because commercial incentives are closely related to avoiding warning signs from a company's models. Note that the mechanism is distinct though: The previous equilibrium is about whether similar safety work would've happened anyway. But even if you managed to make more (commercially useful) safety work happen than would've happened otherwise, the long-run effects of that may still be neutral in expectation because you reduced the probability of warning signs.
Similar to our first equilibrium, risk awareness strongly shifted upward recently through the Hugging Face incident. The way I would frame it, the equilibrium point is continuously rising as AIs become more capable and thereby harder to align, and OpenAI's safety effort didn't rise likewise and was thus below the equilibrium point, making warning signs more likely, which surfaced as the Hugging Face incident.
Considering how safety work may affect the probability of warning signs is now more important than ever, because since the Hugging Face incident it looks a lot more plausible that we'll get further warning signs which may be able to trigger (hopefully good) government action. I think having more warning signs would be great. (As example of what I mean by a (relatively severe) warning sign: Something like the Hugging Face incident, but the AIs also managed to exfiltrate their weights and hacking more datacenters and cryptocurrencies, but the rogue AI is ultimately contained.)
(I'm not sure much changes if existential safety doesn't strongly dominate your concerns; the same equilibrium applies to other problems. If problems are more visible this increases the chance that liability laws will be imposed on AI companies. The next section focuses on optimizing existential safety though.)
How might we want to invest in safety research then?
The risk awareness equilibrium applies only until AIs are able to take over the world, because afterwards inadequate safety likely doesn't materialize as warning sign but as takeover. If you think that by default the level of safety after that point is likely insufficient to prevent takeover, you may hope that until then very large warning signs appear which cause a treaty. From this perspective, safety work that reduces the probability of warning signs may not just be useless (as the risk awareness equilibrium suggests) but even a bit harmful.
Let's consider this simplified model which distinguishes 3 phases of AI capabilities:
- Phase 1: AIs are not capable enough to take over the world even absent AI control measures.
- Phase 2: AIs would be able to take over absent our AI control measures but not with those measures.
- Phase 3: AIs are able to take over. By that point AIs will need to be aligned or corrigible and will likely be smart enough that human efforts no longer matter.
In Phase 2, it is of course important to prevent takeover, and warning signs would then more likely take the form of company-internal incidents that get caught. So from that point on it seems very important to have good control measures in place, whereas before that point good control measures might prevent warning signs from becoming larger and externally visible.
We ideally want as many warning signs to appear as possible while sacrificing as little safety as possible. Let's look at different areas of technical safety research to see what is best there:
- AI Control (and security): Ideally AI control is inactive in Phase 1 (so it doesn't suppress warning signs) and active in Phase 2. I think this means we don't want to implement AI control techniques in companies yet, although I'm not sure how much time control implementation takes and when we should start. AI control research that is kept secret until later sure seems useful. I'm unsure about public control research.
- Misalignment evaluations: Demonstrating problems seems useful for raising risk awareness, although in Phase 1 this could also lead to catching problems early and reducing the chance of strong warning signs. So I'd say it mostly looks good from Phase 2 onward.
- Alignment: Reduces warning sign probability until Phase 3. How much it helps with the probability that AIs are aligned in Phase 3 is a complicated question; I think work on aligning current AIs likely doesn't help that much, but details are out of scope of this post.
- More ambitious alignment approaches which may only be useful later have less of a risk of reducing warning signs.
To be clear, this is just a simplified model. In practice we may have large uncertainty about whether AIs could take over, and weighing the tradeoffs here requires detailed considerations. And of course considerations like whether research might speed up AI capabilities are also important.
Aside from technical research, I think the possibility of strong warning signs makes political advocacy more important, so large warning signs can actually be turned into effective policy.
Conclusion
If you're doing safety work, my recommendation is just to think about your theory of change clearly. If you have a high salary, consider donating.
If you're funding stuff, keep in mind that work in the vicinity of what AI companies are incentivized to do anyway may often not have much impact.
I recommend funding xrisk policy advocacy or work on developing treaty verification mechanisms. E.g. ControlAI, MIRI. It's much more funding-neglected than technical safety anyway.
Appendix: But isn't there also an equilibrium for policy advocacy?
Yes there is! If we funnel a lot of money into policy advocacy, tech companies will likely also increase their lobbying spending.
However, even if we assumed tech companies could just cancel our progress, it would at least cost them money, and probably much much more because we have the asymmetric advantage of truth and people don't trust tech companies. And there are likely diminishing returns to spending more money, so if we invest a lot they may not be able to cancel our progress at all.
And I think it's very important to consider that we may likely see more warning signs. If we get a big warning shot, lobby organizations who were warning about the risks and also have plans ready to be implemented may have a lot of influence. It's quite plausible to me that we would have better pandemic preparedness now if there had been well-funded lobby organizations warning about pandemics and advocating for pandemic preparedness measures before Covid-19 hit.
Thanks to Justis Mills for feedback on this post.
- Or at least I think the counterfactual should be evaluated against non-EAs, though it's debatable.
- Though if you believe that an AI company, e.g. Anthropic, is actually mostly caring about the good of humanity and also sane enough to pursue that in a good way, then saving that company money actually counts as a positive externality.
- I don't actually think this negative externality is that important considering how small the safety budget of AI companies is. Though possible that it will become a bit more important. (And of course some AI safety work has other negative externalities like applications for AI capability that may be more significant.)
- But if it is work that substitutes for AI company work, you bear the small negative externality more fully because you don't have the small positive externality of costing the AI company some money. And in fact if you produce public work you bear the negative externality for all AI companies that use it, not just one! (Though a bit unclear how responsibility is split between you and the person who decided to fund you.)
- Only a bit harmful because safety work for reducing warning signs might happen anyway. E.g. not doing safety work could cause earlier smaller warning signs that then trigger safety work that prevents bigger warning signs anyway.
- This phase corresponds to the post-handoff phase in the alignment roadmap.
- Though aligning frontier AIs may become more useful in the future.
- E.g. I think it's quite useful if politicians don't just hear that there was some incident where an AI hacked a company, but have someone explain it to them in detail. A friend of mine who is doing policy advocacy in Germany mentioned politicians are often sort of shocked when he explains the Hugging Face incident in detail.
- I know a few people doing political advocacy in Europe who I think would be even a bit more cost-effective to fund, feel free to PM me if you're maybe interested in providing funding.
- At least xrisk advocacy. I am not an expert on how funding bottlenecked technical governance research is, but my guess is still significantly.
- Although I'm by no means an expert on US policy, and admittedly there wasn't a strong anti-pandemic-preparedness lobby. But seems quite plausible that the AI lobby will be less listened to if a large warning shot happens.
- I don't know when that point is. It could be soon.