Two Axes of Alignment: A Framework for Robust Superintelligence Alignment

1. Summary

I classify alignment research along two axes: forward-chaining vs. back-chaining reasoning and extrapolative vs. invariant justification of the safety property in question. I argue that extrapolation is insufficient to justify confidence that the safety property will hold while crossing into the superintelligence capability level, whereas an invariant justification is necessary. I also claim that while forward-chaining from current models may give us useful safety properties and even local invariants, back-chaining from superintelligence aims to find the jointly sufficient set of safety properties for alignment. Hence, I argue that robust superintelligence alignment (denoted RoSA instead of RSA, to avoid confusion with RSA encryption) requires approaches that back-chain from superintelligence and establish invariant safety properties. Some of the ideas here draw on existing alignment thinking. My aim is to synthesize these ideas into a useful framework for considering alignment approaches, specify the requirements for robust alignment, and give potential objections.

2. Two Axes of Alignment Research

Axis 1: Research Direction (starting point of reasoning)

Forward Chaining (forwards from current AI): Start with current systems and develop methods that make them safer as capabilities increase.

Backchaining (backwards from superintelligence): Determine properties necessary to align a superintelligent AI, and then work backwards to develop methods or architectures for alignment.

Axis 2: Basis of Reasoning

Extrapolation: Observe a method that preserves a safety property for a model at a capability level, then infer this holds as capabilities scale. Methods that are part invariant still classify under extrapolation, as the justification still must rely on a fundamental extrapolation.

Invariant: A safety property has a formal or structural argument establishing its invariance across capability transitions.

The Four Quadrants of Alignment Research

Axis 1/Axis 2

Extrapolation

Invariant

Forward Chaining

RLHF/RLAIF; behavioral evaluations and red teaming; model-organism experiments; empirical AI control; pragmatic mech interp

Ambitious mech interp

Backchaining

Empirical tests of debate and IDA; weak-to-strong generalization

[RoSA lies in this quadrant] Agent foundations; alignment focused learning theory; theoretical analyses of debate and amplification

Note: The quadrants are idealized categories. In practice, there may be overlap for alignment approaches.

3. Why Extrapolation is Insufficient for RoSA - Persistence

Let an alignment target specify how the model is to act aligned. I will remain agnostic as to the correct alignment target. For this framework a model will be considered robustly aligned if it satisfies some specified alignment target at some reliability, independent of domain.

Our confidence in alignment can be justified via an extrapolative or invariant basis. Empirical success forms an extrapolative basis. A formal or structural proof forms an invariant basis.

Thus, I can state my argument as follows:

Premise 1: Superintelligence will operate outside of the capability regimes in which our alignment methods can be empirically tested.

Premise 2: Empirical success within one regime does not establish that a safety property will hold outside of it, especially as novel failure modes may arise in new regimes.

Therefore: Empirical success of a safety property within a pre-superintelligent regime does not guarantee that property will hold in a superintelligent regime.

A superintelligence is robustly aligned if we can provide an invariant formal guarantee that the alignment target is satisfied in all viable environments. Any claims of alignment that rely on extrapolation of empirical evidence for justification are non-robust.

Thus, we see there is an issue of persistence. We cannot guarantee a safety property will persist outside of regimes it was shown to have empirical success in.

4. Why Forward Chaining is Insufficient for RoSA - Completeness

Following from establishing the need for invariants to determine RoSA, we can then similarly address the second axis of forward-chaining versus back-chaining and have a quadrant that specifies approaches to determine RoSA.

Forward chaining from current models may be instrumental to RoSA (refer to section 5) even if FC does not itself constitute a robust solution to superintelligence alignment.

The problem is one of completeness. Forward chaining from current models may allow us to discover local invariants that hold across regimes. However, because FC is bottlenecked by models we currently have, we may overlook potential failure modes where invariants are needed when in the superintelligence regime. Solving every failure mode of today with an invariant does not guarantee we have the required set of invariants for superintelligence. Essentially, FC may discover some/all of the required invariants, but it can’t establish the set is complete.

The required set of invariants can only come from back-chaining as this would require inspecting the potential failure modes of the regime that we expect to deploy our model in. If researchers are able to find all the failure modes, we could construct the jointly sufficient set of invariants to guarantee RoSA.

Thus, I can state my argument as follows:

Premise 1: RoSA requires a set of safety properties jointly sufficient to satisfy the alignment target in the superintelligence regime.

Premise 2: Forward chaining does not guarantee that the failure modes observed will establish the properties jointly sufficient to satisfy the alignment target in the superintelligence regime.

Premise 3: Establishing completeness requires backchaining from the alignment target in the superintelligence regime to determine safety properties jointly sufficient to satisfy the target.

Therefore: RoSA must include backchaining from superintelligence to determine the complete, jointly sufficient set of invariants.

A superintelligence is robustly aligned if we can provide an invariant formal guarantee that the jointly sufficient set of invariants satisfies the alignment target in all viable environments. Any claims of alignment that rely on extrapolation of empirical evidence to justify or derive a set of invariants from forward chaining are non-robust.

5. What Forward Chaining Is Good For

The ways forward chaining is valuable:

  1. Bootstrapping robust alignment: Forward chaining may be used to develop an automated alignment model strong enough to provide backchained safety properties and invariant justification, which (hopefully) we can then independently verify.
  2. Reducing risks from current and near-future models: Evals, red teaming, control, and preference fine-tuning can mitigate failure modes in current models.
  3. Generating empirical information about alignment failures: Finding failure modes and seeing how alignment properties fail as capabilities increase.
  4. Providing testbeds for theoretical ideas: These tests may lead to invariant or structural justifications down the line.
  5. Enabling a safer transition: These methods may buy time for RoSA.

Thus, I do not claim that forward chaining has no value, as FC may contribute substantially to reaching RoSA. My claim is instead that FC alone cannot establish RoSA.

6. Conclusions

The central claim of this post is that we need backchaining for completeness and invariant justification for persistence to establish RoSA. Forward chaining is clearly useful in the buildup of this idealized alignment solution, particularly in the case of automated alignment, but FC itself will not establish RoSA.

Thus, FC and extrapolation may help us reach RoSA, but backchaining and invariant justification are required to establish RoSA.

7. Objections

Extrapolation is better than I claim.

Perhaps empirical evidence across sufficiently diverse systems can provide extremely high confidence. An alignment method might succeed across sufficiently diverse regimes to justify high confidence, even without an explicit invariant. Simply stating “extrapolation is not proof” may be insufficient justification.

Robust superintelligence alignment may be impossible.

Establishing the alignment of a superintelligent model with the required confidence may be impossible. Backchaining may be fundamentally impossible if we cannot find a useful framework to consider a superintelligent model.

Thank you to Alec Harris for useful feedback and discussion. Thank you to my girlfriend Sonya for proofreading.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论