Minimum required pacing: finish aligning it before you train the next one
Let's say we coordinate to pace the frontier, but we don't coordinate on an outright pause. What should the standard of pacing be?
My proposal for a minimum requirement: a lab cannot train the next generation of models (in terms of capabilities) until they have adequately resolved all significant alignment issues in the present generation of models.
You don't launch a jet that goes at Mach 3 before you've figured out why your Mach 2 jet keeps crashing. With any other kind of dangerous technology, we wouldn't accept "shrug and move on".
This proposal may appeal to many prosaic-alignment optimists and pessimists, because they usually differ on how easy it would be to fix current alignment issues with current techniques.
- Prosaic-alignment optimists might expect that a moderate coordinated slowdown would suffice to adequately resolve alignment issues at each generation.
- Prosaic-alignment pessimists might expect it to result in an extended pause, buying time to discover new fundamental approaches to aligned AI.
What might be adequate?
What might it take to adequately resolve alignment concerns for a generation of models? Here's my vague, qualitative proposal:
- Multiple independent alignment organizations examine and red-team the newest model at many checkpoints throughout training, with full access to the data and with the best tools available.
- Alignment issues need to be taken seriously whether there's a metric for it or whether something is qualitatively off, whether the model has a bad attribute or whether they can't robustly conclude it has a good one, etc.
- Alignment issues early in training may not be a showstopper, but there must be a robust safety case for why the later stages of training will resolve those issues.
- A model of a new generation is ready for restricted deployment only when the alignment organizations are satisfied with it.
- Restricted deployment starts with a group of trusted users doing tasks that don't contribute to model training and monitoring. The alignment organizations monitor restricted deployment for any new issues.
- If any significant alignment issue arises during restricted deployment:
- Pull the model.
- Do a root cause analysis of the issue and why it was not caught during training.
- Figure out a principled way to fix the disease rather than the symptom.
- Train the next model of the current generation accordingly, and start again at the top.
- A model of a new generation is ready for deployment (internal or external) only when it's survived a significant period of restricted deployment without any issues arising.
- If any significant alignment issue arises during broader deployment, hit reset as outlined in the previous section.
- If no alignment issues have come up over a significant stretch of deployment time, only then can the lab begin training a next-generation model.
I still wouldn't say that this is fully adequate, but I'd treat it as a minimum requirement.
This proposal is incomplete; you can help by expanding it
This isn't a concrete proposal, just a heuristic one. A few obvious missing pieces, aside from the standard implementation problems:
- How do we define "a generation of models", within one lab or across several labs?
- What should we do if an intervention to fix an alignment issue also significantly increases capabilities?
- How would we define and enforce an agreement, if we're unable to steer solely by predefined metrics?
- When is an alignment issue big enough that we should halt the model? There will be obvious pressure from the lab to set the threshold high.
- And no, "scores better than previous models on our metrics" is not adequate. If a user points out a weird but alarming thing that wasn't noticed before deployment, that is a critical failure in the process as well as in the model.
- Sure, maybe your Mach 2 jet crashes for a reason that stops being a problem at Mach 3. But it would be crazy to assume that!