Almost nobody is funded to figure out what work would solve alignment
Solving alignment would be easier if we worked out what problems we actually need to solve. This could be called the alignment meta-problem. Work on this problem is rarely directly funded. More focused work on it should let us use our limited time and funding more efficiently.
If we're driving toward a cliff, maybe we should buy better headlights.
All too often we're doing work that merely sounds or feels good, and optimizing less than we could for work that drives most efficiently toward success. Some of this is inevitable and some of it is useful, but we could do more to light the path ahead.
Most researchers agree that mech interp, refining and improving alignment training, control, theory, and miscellaneous techniques like confession are useful for solving alignment. Working toward regulation and slowdown/pause is also commonly considered useful in the governance space, and spreading awareness of alignment risks is pretty obviously useful for accelerating and enabling all of this work. But we don't know what variants of this or other work make best use of our limited time and funding. Which of them, or what others, are the most efficient use of resources? I think that we could probably get usefully better answers at relatively low cost.
Currently work on the meta-problem is done either in researchers' spare time, during grant applications, or inside funding orgs and AI developers. All have limited incentives to publish their thinking legibly, and incentives during grant applications and inside dev orgs are subject to Goodharting and motivated reasoning: sounding good vs. being good. More funding directly for analysis and planning seems like an efficient leverage point.
The AI Futures Project is the exception that demonstrates the trend. They are funded to do prediction work, but that spreads into many aspects of the meta-problem. They make Gears-Level Models of how alignment and governance work might succeed or fail, and so identify work that's likely to be more and less useful. They have the time and focus to explicitly include the many cross-dependencies between governance, public opinion, and technical alignment work, in contrast to people who do this in their spare time and mostly focus on their own area of expertise. (Other orgs and people do aspects of this work, but arguably do less direct work on the meta-problem as I'm thinking of it.)
Many researchers and orgs do some of this work despite it not being directly funded, because they think it's so valuable. And people at funding orgs do meta-problem work as part of their job. But publications of funders' models tend to be short and limited (although kudos to the funders who do carefully work out and publish their theories-of-change!).
Why not to fund more work on the meta-problem
There are good reasons not to fund this type of work more heavily. In brief:
- Science doesn't typically do this, for good reason
- Could concentrate funding, neglecting important work
- Bad predictions and work recommendations are worse than none
- This work is being done adequately through indirect incentives
- Maybe what we're doing is the right work, particularly for short timelines
- Could waste money on useless or bad work that's hard to evaluate
- Prediction is hard, particularly about complex and novel scenarios
These are all important points. Each is something to guard against, or a reason to not over-fund the meta-problem.
Arguments in favor, compressed
Each of the arguments against suggests an argument for:
- Unlike other sciences, alignment has a specific goal and time pressure
- Like planned projects such as Apollo, Manhattan, WWII
- (And perhaps other sciences would benefit, too; observing cognitive neuroscience and its funding for 23 years has led me to think they would)
- Over-concentration of funding is more likely if there is less/worse meta-problem work
- Over-concentration could be happening now, since only a few people and orgs are funded to do this work
- We're already making predictions and work recommendations, just rarely with much time-on-task
- Could draw more funding into the field by clarifying what's probably useful
- While a lot of this work is being done, most people are doing it in small pieces, and are incentivized and have Motivated reasoning toward justifying their own work
- Maybe we're already doing all the necessary work; why not rigorously check?
- Because it's hard to evaluate this work, streetlighting steers away from it
- It's also arguably an Illegible problem in Wei Dai's usage
- Prediction is very hard, but also very useful.
- Actually funding people to do it and publish it full-time might be better than everyone just doing little pieces in their spare time
This post is a teaser for a longer, more worked-out draft that's been stuck in my queue. If responses here change my mind or indicate that No One Cares, it will probably stay unfinished. This is pretty telegraphic; I can elaborate, give examples, and clarify in the comments.
All questions, feedback and suggestions are appreciated!
Disclosure: Much of My research has recently been on aspects of the meta-problem. This is a result of increasingly thinking such work is even more neglected than the object-level technical alignment work I started out doing. I feel very lucky to be able to devote time to this type of work.
Disclaimer: I'm not deeply familiar with all the orgs, and I'm sure I've overlooked orgs funding overlapping work. Apologies to those I've missed.
Thanks to Peter Gebauer and others for many conversations touching on this claim, and Peter for comments on an earlier draft. This framing was inspired when a donor friend said "I might fund more alignment work if I could tell what would be useful," causing me to say "Maybe we should be funding figuring that out!"