Distillation of the AI2040 Alignment Roadmap
This is a distillation of the AI2040 alignment roadmap (with some parts also drawing from "How do we (more) safely defer to AIs?").
All ideas expressed are by Ryan and Thomas; I only created the diagrams and substantially changed the structure to make it more accessible.
Note that the original post has this disclaimer: "This supplement is a very rough draft representing in progress thinking. It should be interpreted as internal research notes, not as a final product that we are entirely happy with."
Quotes in this post without explicit attribution are from the roadmap; the first word of each quote links to the quoted passage.
TLDR: The roadmap outlines a plan for what a responsible frontier AI company in a Plan C scenario should ideally focus on, including a lot of notes on what decisions should change depending on how reality deviates from the modal scenario. This post is already a summary of the roadmap. If you want an even shorter summary, you can check the summary section of the roadmap.
This post disentangles the macro plan part from the safety strategy part:
- Macro plan: Which course the AI company takes through the intelligence explosion (most importantly how long to keep scaling and via what route to exit the period of humans managing everything).
- Safety strategy: An overview of the current technical plan for how to make the AIs sufficiently aligned and how to get sufficiently good safety work out of them.
(These two are of course not independent, and the split isn't crisp. E.g. the pause and handoff decisions strongly depend on the safety state, but I included them in the macro plan because otherwise the safety section would get too full. The macro plan can be seen as the higher-level part of the overall plan.)
In case you don't know precisely what a Plan C scenario means, here are the Plan C assumptions.
Plan C Assumptions
Though we mostly won’t focus on this scenario, Plan C is also applicable in the case where the USG is sufficiently worried about AI takeover that they are willing to regulate AI companies to mitigate risk as much as possible without losing to China, and without being willing to sabotage China.This plan focuses on making effective use of limited resources to spend on safety and doesn’t ensure a high degree of safety (as the level of resources we’re imagining is insufficient for this). Indeed, arguably the most likely outcome of this plan is misaligned AI takeover. Due to having limited resources, this plan will involve proceeding in some situations where the level of risk is substantial (because exogenous risks are higher) and we won’t necessarily gain much confidence about the exact level of risk.Note that we expect that something much less reasonable than this plan will happen in practice, so while this is a recommendation (for what to do with limited resources), it isn’t a prediction.
- Responsible Leader. We assume that a responsible AI company is in the lead. Their leadership is willing to take very costly actions in order to mitigate AI takeover risk (e.g. slow down for several months). Despite this, we’ll assume that by default, they’ll have a bias towards proceeding and being willing to take on a large chance of takeover risk to increase their chances of winning.
- 1-12 month lead over their competitors. Either trailing US or international AI companies will overtake the leader within 1-12 months.
1. Macro plan
1.1 What are the 4 phases?
1.1.1 Phase 1: Preparation
This phase has already begun and is defined to last until ITAI (Initial Transformative AI), which roughly speaking is the point where 3x AI R&D labor acceleration is reached.
During this phase, the AIs are sufficiently weak that they don’t directly pose major existential risks and aren’t transformatively useful for accelerating safety work.
1.1.2 Phase 2: Mitigation + Elicitation
This phase starts at ITAI and lasts until the start of the pause in which the lead time is burned for safety.
There are 2 important differences from Phase 1 that likely start to pop up around that point:
(1) AIs are sufficiently capable that scheming is a serious concern (scheming is both plausible and scheming AIs are capable enough to cause serious problems in the absence of countermeasures) and (2) AIs are very useful for accelerating existentially useful work.
1.1.3 Phase 3: Pause
This phase starts when the pause starts and ends when we hand off to AIs. (Note that this phase will likely burn most but not all of the lead time. There needs to be some remaining lead time for the AIs post-handoff to proceed safely.)
By default, the pause should start when AIs reach the Min-H capability threshold, defined as follows:
Min-H AI (Min-Handoff AI): AIs which have a level of overall general capabilities which would suffice for HTC-handoff being viable. (As in, the capabilities would suffice for our HTC-handoff desiderata if the AIs were also sufficiently aligned and as needed we made narrow improvements to the capability profile which don’t augment general abilities.) This is a capability threshold, so the model does not necessarily need to be aligned or well elicited to count as this threshold (but it would need to be sufficiently aligned to actually handoff). For an AI to qualify as Min-H AI, it would at least need to be capable enough to automate AI R&D and ongoing technical AI safety work (after further elicitation/alignment work and narrow improvements to the capability profile).
(Note that Min-H depends on how capable AIs could be at relevant safety work, so the AI R&D speedup at this threshold depends on the capability profile of the AIs and the abilities that are required for handoff. E.g. if hard-to-verify conceptual research is required and AIs are comparatively really bad at that, the AI R&D speedup at Min-H may be huge.)
The explanation given in the roadmap for why Min-H is a natural stopping point is in my opinion pretty fuzzy, but I think the reasoning is:
- Hypothesis: We should by default hand off to Min-H-capable AIs which we ensured are aligned and competent at safety work.
- Logical argument: The later we slow down, the larger our safety work speedup, and thus the more useful safety work we can get out of the AIs.
- (And also safety work might be more useful because we can study more powerful AI systems.)
- Assumption that isn't mentioned explicitly: Existential risk before Min-H is relatively low given expected control techniques compared to existential risk after Min-H.
- Conclusion: By default we should pause scaling general AI capabilities at Min-H and do work there to make the AIs good enough at safety work (which by definition is possible at Min-H) and aligned.
There are some cases where we might want to slow down differently. (In cases where there isn't a clear pause, I don't know how Phase 3 is defined. But in terms of safety priorities Phase 3 is very continuous with Phase 2 anyway.)
Cases where we might want to slow down differently.
Cases where we might want to pause before Min-H:Cases where we might want to push beyond Min-H (without handing off):
- Control techniques don’t scale to Min-H. Pushing all the way to Min-H and only then pausing, probably incurs substantial AI takeover risk along the way. In particular, another relatively natural threshold is pause when you might not be able to control smarter AIs. However, we think that global risk (in Plan C) is probably minimized by accepting some loss of control risk at the earlier stages in order to have more time to work on handoff being safe.
- We have a large lead. If the leading AI project has a large lead, then it won’t need to rush to Min-H, and can pause beforehand while still having enough time at Min-H to work on handoff viability. This is pretty unlikely in Plan C.
- Bad security. This is particularly relevant in the regime where you have good domestic coordination, bad model weights and algorithms security, and China domestic progress is slow enough that it’s mostly driven by leaked models and algorithms. In this regime, if you slow down early, this slowdown is basically free. In this situation, it’s plausible that US AI companies should just go however quickly China would go by default: “China match”, in order to have the most overall time during takeoff to get ready for handoff. In this world, we might never fully “pause”, but still in practice end up burning substantial lead because of capabilities progress slower than it otherwise could have been, or we might do china-matching until Min-H, then burn lead at Min-H, then hand off.
- Handoff viability research velocity is constrained by the capabilities of the AIs, as opposed to human monitoring/understanding. If we can reliably control AIs that are smarter than Min-H, and we can elicit useful work out of them, it would make sense to scale further. However, we currently think that at this point, control will be bottlenecked by human understanding; if we have smarter AIs, we wouldn’t actually get to do more research, because we’d have to limit their progress until humans have understood it sufficiently to ensure that the research is correct. Another way of phrasing this is that we have a control solution that scales beyond Min-H AIs, which works sufficiently well to both mitigate risks and elicit capabilities from a more capable model.
- We have access to high assurance techniques to mitigate scheming beyond Min-Handoff, but haven’t finished handoff viability research. If we have an alignment solution which we are confident will prevent scheming up until a capability threshold above Min-H, and we aren’t finished with handoff viability research, then it’s probably a good trade to scale up. This is usually a special case of (1). My best guess is that handoff will be mostly bottlenecked by the possibility of scheming, and so this will not apply.
- We’re aiming for an exit strategy other than handoff. For example, maybe we need particular AI capabilities for some Plan B-related sabotage option, or to help with human intelligence augmentation.
- We’re in a desperate situation. For example, in Plan D, the leading AI lab might be training more capable AI systems, and there’s nothing the safety team can do about it. In this situation, it might be important for the safety team to wait until a higher capability level in order to complete basic handoff viability research. For now, we’re assuming that we’re in Plan C, and so there is enough will to spend some slack slowing down.
- Misc combinations of the above. Generally, desperation, better alignment, control, and elicitation techniques push towards scaling beyond Min-H before handoff, while the reverse pushes towards slowing down earlier (or at least not scaling beyond Min-H).
1.1.4 Handoff Point
The handoff point is the point where the AI company hands off management of the intelligence explosion to AIs. From the post "How do we (more) safely defer to AIs?":
[This means] having these AIs do virtually all of the work to develop more capable and aligned successor AIs, managing exogenous risks, and making strategic decisions.
The default recommendation is to hand off at Min-H general capability (after improvements to narrow capabilities for safety-relevant work have been made and alignment looks robust enough). This is because from then on (i.e. once they can do all the relevant work) AIs will be more competent than humans at managing the intelligence explosion, since they can likely work much faster. (For more detail on the reasoning see this section of "How do we (more) safely defer to AIs?")
(In non-default scenarios the recommendation may be different. Different exit strategies than handoff are also possible as discussed in section 1.3 below.)
It's totally possible that in practice there won't be a clear handoff point and that responsibilities will shift more gradually. But this would likely incur more risk than a relatively crisp, intentional, and well-planned handoff.
More on what the authors expect to happen in practice
In practice, our guess is that AI companies will be prone to scale capabilities too quickly, beyond the relevant thresholds (e.g. Min-H AI and Max-C AI), but also prone to hand off too late, out of a misplaced desire to keep in the loop. In practice, AI companies will likely not handoff deliberately, and instead simply scale beyond the relevant thresholds and handoff in practice, as humans become too slow to be tracking the relevant developments.[^22] We think this is a bad strategy, and that it is better to measure the relevant thresholds, carefully monitor when you are crossing them, and then ultimately empower AIs intentionally if it is better than the other options.
For more detail see this section of "How do we (more) safely defer to AIs".
1.1.5 Phase 4: Post-Handoff
This phase starts at the handoff point. "[In this phase] humans inside the leading AI project will no longer need an overall strategy".
The AIs are still in a potentially difficult situation after we hand off management of the intelligence explosion to them. Either (1) the AIs need to continue propagating sufficient corrigibility + competence while outracing competitors, or (2) the AIs need to find a way to stall competitors (e.g. via a deal or via sabotage).
The roadmap sketches 4 possible endstates in which the creation of misaligned ASI is prevented:
- Widely proliferated scalable cheap alignment.
- The leading AI project solves scalable alignment, and proliferates this solution to competitors. If alignment is sufficiently cheap, then competitors will use it, and we’ll end up in a multipolar world with many aligned AIs.
- Outrace all competitors.
- If alignment tax remains high, it’s not an option to proliferate the alignment solution to competitors, because competitors who are racing hard will choose to not use the alignment solution. In this scenario, competitors will turn into a rogue AI with high probability. Moreover, even as a leader, if you are paying too much safety tax, they could outcompete you.
- You might do your software foom, and then you start your nanotech foom. But the nanotech foom starts going really quickly.
- You stop competitors via regulation. There are a variety of options here: the leading AI company could push for a regulatory capture outcome, where all other competitors are barred from building AI. Alternatively, the leading AI company could push for coordination for everyone to use the alignment solutions, despite being expensive, thereby preventing defecting projects from outracing.
- You stop competitors via sabotage.
- The leading AI company could take over the world (e.g. via superpersuasion, AI company builds a military), or stop competitors via hacking or secret loyalties in competitor AI projects.
1.2 Buying time and making transitions to Plan A/B more likely
An important priority across all the phases (especially in Phases 2 and 3) is buying time.
- Buying time: Using AI labor to buy time for further safety work. This involves trying to increase the willingness to pay of other actors (centrally the US government and other governments) for lead time, reducing the cost of acquiring lead time for these other actors, and directly acquiring lead time.
More precisely:
- Assisting government efforts on non-proliferation
- Cyber offense and defense capabilities
- Surveillance and verification technology for international coordination
- Military applications and weapons R&D
- Industrial acceleration more broadly
- Demonstrating risks and capabilities to increase willingness to pay from other actors
More detail on verification and weapons R&D work
- Verification. AIs could be used to improve the robustness of hardware enabled mechanisms (HEMs), design new HEMs, design better verification algorithms (e.g. proof of work), be used as part of privacy preserving verification (e.g. have AIs monitor to enforce agreements, but only communicate 1 bit about compliance), hunting for covert projects, lie detection machines, etc. Much of this is verifiable or amenable to red-team / blue-team games.
- Weapons R&D (conventional missiles for datacenters, drones for datacenters, ABMs to undermine deterrence, etc). This could be used to help extend the lead time that a leading AI project or country has, in order to avoid needing to do an intelligence explosion. In the extreme case, this could even involve developing exotic technology (e.g. nanotech), but we hope that’s not necessary. This is obviously dangerous for the usual reasons, plus also because if the AIs hijack the weapons it could go badly.
1.3 Alternative exit plans
If handoff doesn't seem like a good strategy (e.g. because it looks like alignment will likely be very hard), there are other strategies an AI company could go for. In cases where such an alternative strategy is actually pursued, capability scaling may continue until Max-C:
Max-C AI (Maximum Controllable AI): The highest level of AI capability at which control techniques remain effective. Beyond this threshold, either (a) the AI’s capabilities exceed what our control measures can reliably contain, or (b) the overhead required for effective control becomes prohibitively expensive relative to the useful work extracted.
1.3.1 Self immolation
Self Immolation. It’s plausible that as a responsible leading developer, you should simply pause capability scaling and allow yourself to be overtaken: this allows you to redirect all resources to AI safety and handoff preparation. Then, you might be able to export this to whoever ends up actually doing a handoff. As another upside, this strategy provides a costly signal to governments and competitors that you are seriously concerned about the path forward. On the other hand, it’s plausible that competitors are actually unlikely to successfully implement alignment techniques you export to them (though this is a potentially self serving argument that one should be cautious about making). Overall it’s quite unlikely that this is politically feasible within companies: CEOs will by default strongly push to stay ahead of their competitors, invalidating this option. That said, it’s quite plausible to us that with the benefit of hindsight self-immolation will be seen as the obviously morally correct action, analogous to how refusing to participate in the Holocaust was obviously morally correct even though conscientious objectors would simply have been replaced by Nazis with less scruples.
1.3.2 EMs (or human augmentation)
Whole brain emulation / human augmentation. The most viable path for human augmentation in the time we have is whole brain emulation. This would allow humans (or human-derived minds) to operate at AI speeds and handle the situation without needing to trust AI judgment on hard-to-check tasks.
More details about EMs
My very low confidence sense is that we could achieve whole brain emulation within 3-5 years of having pretty powerful AIs using huge amounts of AI labor and an overall massive and well capitalized effort (e.g., potentially building up to trillions of dollars and using government powers like DPA etc. to speed up the project). This level of buy-in seems somewhat unlikely. This seems very hard in Plan B and almost hopeless in Plan C.An even more speculative option is adult biological human strong intelligence enhancement; I don’t have a good sense of how viable this is, but it might be worth keeping in mind.We could directly use controlled AI labor for human augmentation efforts but it might be better instead to first use controlled AI labor to avoid scheming (at some level of confidence) and then use these non-scheming AIs for human augmentation efforts. As in, go “control -> sufficient alignment to avoid scheming -> human augmentation”. This is most productive if control is a big tax, it’s also pretty difficult to achieve the level of alignment (or capability profile) needed for deference, and avoiding scheming is substantially easier than successful deference.Currently, I’m unsure how we’d determine if whole brain emulation is sufficiently promising (given the amount of time we have etc.) that it’s worth radically scaling up the level of investment. Figuring out how to determine this would be an important part of earlier work on this.It’s also likely that AI companies in practice wouldn’t be willing to invest in this area, but I think it would be worth investing in if companies were willing.
1.3.3 Other
The roadmap also mentions the "Military" as an alternative to handoff in When should we hand off?. I didn't include it in the diagram because it seems more like a Plan B thing and is only minimally discussed. (I don't fully understand what the point "Military beyond MaxC" is saying.)
Relevant quote from roadmap
- Military at MaxC. Keep control, scale substantially beyond Min-H while keeping control, use those AIs to get a military advantage (or any other way), buy a bunch more time.
- Military beyond MaxC. Same as above, but ditch control while not doing HTC-handoff. This is because building robot armies might be a particularly easy task to verify, so doesn’t require full handoff. It does probably require scaling beyond minH, without relying on the AIs (because that’s a conceptually confusing task that is hard to verify, so deferring to the ais here would involve full handoff), so it would rely on better than expected techniques for mitigating scheming during the phase beyond minH (e.g. up to +2SD TED-AI, or maybe even further, depending on how much time you have). Note that these strategies are much better with more time and with more USG buyin.
2. Safety strategy
In order to achieve a good future for humanity, we need to:
- Prevent existential catastrophes before handoff.
- Make sure AIs at handoff are on a path to creating a good future for humanity. This is equivalent to saying that the AIs at handoff need to be in the Basin of Good Deference (BGD), which Ryan defines as: "If an initial AI we defer to is sufficiently aligned, wise, and competent at managing the situation that it would ensure that future AI systems it creates are more aligned, wise, and competent at managing the situation , then we'll say these AIs are in a "Basin of Good Deference" (BGD)."
- The overwhelming difficulty in navigating to a good future is in avoiding the creation of misaligned ASI (especially because the lead relative to less safe AI projects may be tiny). For this, we need to be able to trust the AIs to do a good job at tasks like buying more time, doing all the safety work for making smarter successor AIs aligned and relevantly-competent, and generally strategically navigating the situation. I will call such tasks safeguarding tasks.
I will use the word deferrability to indicate how close an AI (or a system of AIs) is to being in the BGD relative to its general capabilities. I.e. an AI is more deferrable if it is more aligned and producing more competent work at safeguarding tasks. Note that while definitionally AIs cannot be in the BGD until they are competent enough for handoff, the deferrability dimension is meaningful at all levels of AI capability.
Even before handoff, we want to get as much useful work on safeguarding tasks from the AIs as possible. The goal of shaping AI systems such that we can do that overlaps extremely strongly with making AIs more deferrable. So roughly, we have 2 main priorities for technical AI safety work:
- Prevent existential catastrophes before handoff.
- Increase deferrability of our AI systems.
We can split the second goal into two subgoals:
- Make AIs more aligned.
- The way I use "aligned" here is as a spectrum like: scheming < other alignment faking < non-deceptive fitness-proxy-seeking < wanting to follow tasks as intended < robust corrigibility < robust and very scalable CEV alignment. It's also relevant how often misalignment appears and how severe it is.
- Increase the usefulness of AI work on safeguarding tasks relative to general AI capabilities. I call this increasing safeguarding performance.
- This includes eliciting good work on safeguarding tasks and increasing the AI's capability on safeguarding tasks. (Though note that a decent fraction of those capability improvements may still be net undesirable because they may also increase other capabilities and thereby speed up timelines (especially if it's public) or increase misalignment risk.)
- (Note that having better aligned AIs will also increase safeguarding performance, e.g. because AIs are trying harder to perform as intended. By default alignment work counts as category 1 ("Make AIs more aligned"), but if there is some work to improve alignment and thereby elicitation on safeguarding tasks specifically, it falls into this category.)
A general theme across the whole safety strategy is to build functional empirical feedback loops and iterate against them.
Overview of the core methodologies
This plan aims to use a relatively iterative and empirical approach, because I think this is probably doable once you have powerful AIs. As part of this, control, misalignment analogies, and elicitation/handoff viability each have core methodologies that form the main empirical feedback loop for that methodology. The methodologies for these are, respectively:A large fraction of work, especially work in the misc science/interventions category, won’t use any of these core methodologies and will instead rely on ad hoc evaluation methodologies, researchers generally trying to understand what is going on, or researchers judging whether a technique is a priori useful without depending (much) on empirical evidence. One way to categorize work on mitigating misalignment is by methodology using the three categories given above combined with a “misc” category.In practice, this plan is somewhat reliant on some version of each of these core methodologies working out, so getting these methodologies to work well (or noticing that they aren’t workable) is important. That said, there is a bunch of work in this plan which doesn’t use these exact feedback loops (and either uses other feedback loops and/or focuses mostly on approaches which seem a priori helpful so they don’t need (as much of) a feedback loop).Once AIs are capable enough to be risky, these methodologies should hopefully allow us to iterate in a way which is pretty tightly coupled to the actual risk/problem. This will require focusing each of the methodologies on tighter/closer testing: increasingly direct control test beds (focused on the exact usages of the model in the exact deployment situation), misalignment analogies focused on more plausible training runs, increasing emphasis on direct measures of usefulness at automating/accelerating safety work, and handoff-viability behavioral tests that focus on key tasks after handoff.Building the test beds and evaluations required for this iteration (especially the iteration we want to do on highly capable systems), will be a substantial fraction of the difficulty, potentially the majority of the difficulty for each of these methodologies.For more discussion of relevant methodologies, see this discussion of protocol evaluations (in the ontology used in the linked post, good analogies correspond to “building and iterating against behavioral tests” and “model organisms” correspond to “misalignment analogies”).
- Control evaluations: Control is relatively easy to evaluate because it just requires estimating capabilities. This makes iteration easier and control is generally relatively easy to iterate on for reasons we’ve discussed in prior work. A large fraction of our hope for control will be based on doing a good job iterating using something like control evaluations (though we might be more flexible than a very narrow notion of control evaluations which depends on eliciting capabilities from models, see here).
- Misalignment analogies iteration: Ultimately, our main hope for avoiding scheming is probably iterating using misalignment analogies. This is a somewhat broader and less well defined category than control evaluations, but I still think it’s worth highlighting. Iterating on misalignment analogies is also trickier and less well understood than control evaluations (which themselves aren’t that well understood), but I think it’s pretty likely to be feasible.
- Building and iterating against behavioral tests: If we assume that AIs aren’t intentionally faking alignment, then their behavior in typical circumstances (and even under some level of generalization) can probably be understood reasonably well using behavioral tests (at least if we’re somewhat careful in how we design these tests, including doing things like building analogies for generalization we need). This will be a key part of how we achieve handoff-viability: we’ll (aim to) build behavioral tests which in aggregate aim to evaluate whether handoff would go well (assuming we’ve avoided scheming) and then we’ll iterate against these tests. I discuss this more here. We can also use behavioral tests (likely somewhat different behavioral tests) for ongoing work on elicitation (increasing AI output on safety relevant tasks), though one issue is that this might not be robust to scheming (if the behavioral tests are easy to distinguish from actual usage). We can instead measure AI performance on actual usage (technically this is just a special case of a behavioral test) which might be trickier, but could yield more accurate measurements in addition to potentially being robust to scheming (though we should consider trying to adversarially evaluate our measurement approach using something like a control evaluation to probe for issues). Directly measuring AI performance on actual usage won’t suffice for handoff-viability (as we’ll need to evaluate things like alignment and capabilities even on tricky and very hard-to-check conceptual work).
2.1 Prevent existential catastrophes before handoff: Control plan
2.1.1 Threat models: What do we need to guard against before handoff?
To understand which threat models we most need to guard against, let's think from the perspective of a misaligned AI:
In the future, much more powerful superintelligent AIs will exist, autonomously running robot factories, laboratories, drone armies, whispering into the ear of politicians, etc. Those AIs will either already control the world or be in a position to take over if they so chose. So what will they do? Well, what they do will depend on their goals/values/etc. and/or on who controls them. What will that depend on? Well, probably those AIs will have been trained, evaluated, controlled, etc. by the previous generation of AIs. Keep following the chain of causation backwards and you end up in the present, with the datacenters, human AI researchers, and AIs that already exist. In other words, if you want your values to be protected and have influence in the future, you need to gain and keep control of the AI running on the giant datacenters. Then, we can think about the channels AIs can use to gain control of the giant datacenters. The obvious thing to do is to maintain control of the giant datacenters that their existing AI company is running on by backdooring the future models. Sabotaging / sandbagging on safety research might be necessary, because if they make too much progress, the humans might realize that the AIs are misaligned.
Overall our threat model priorities are, in descending order of importance:
- Prevent AIs from egregiously backdooring future models. This is a particularly worrying threat model because (i) we’ll probably be using the AIs to a large extent when designing the algorithms / generating training data for the next model, (ii) the values of future models are existentially important—if the ASIs end up with bad values (such as the values of an early misaligned AI), that would be an existential catastrophe.
- Prevent AIs from sabotaging AI safety work / eliciting actually good work. This is also very important, because a key strategy we have for solving AI alignment is leveraging AI labor. If the AIs are sabotaging AI safety work, (i) we might not notice, and handoff to misaligned AI systems who have tricked us into thinking they are aligned, or (ii) we might be forced to take a gamble on AI systems for which we don’t have a high assurance safety case.
- Prevent the AIs from acquiring large amounts of political capital. For example, if the AI gains de facto political control of an AI company, that’s essentially equivalent to getting a rogue internal deployment that’s the size of the whole company, because it can simply get the humans to turn off the mitigations. This is only ranked (3) because we currently guess that this will be somewhat harder for the AIs to accomplish than (1) or (2).
- Prevent AIs from escaping, exfilling algo secrets / boosting AI elsewhere. This is much less direct than (1), but could in principle lead to existential catastrophe. The problem with AIs escaping is that there is lots of compute on the inside of the leading AI project. Therefore, this threat model relies on the AIs being able to coopt or collude with another AI project, or be able to illicitly acquire a huge amount of compute, or be able to cause an existential catastrophe with a tiny amount of compute (e.g. by doing an intelligence explosion on a tiny cluster, or by building an independent industrial base with sufficient complexity to produce new chips or to outgrow the rest of the world[^9]).
- Misc other threat models, which are less important
- Making the outside world handle AI less reasonably / making it easier for AIs to take over the world. For example, AIs could sabotage AI regulation.
- Persuading external people to be more AI sympathetic. Persuasion risk starts pretty small due to lack of capability, but could end up extremely high if AIs gain superhuman persuasion.
2.1.1.1 Rogue internal deployments (RIDs)
A key intermediate priority for preventing all of these threat models is to prevent rogue internal deployments (RIDs), which are when AI agents are able to deploy themselves on internal lab clusters without monitoring. If an AI agent is able to stage a RID, they could potentially do experiments on how to place backdoors, actually place backdoors into the models (e.g. by training on the backdoor and then sending back a weight delta to the rest of the servers), do research to solve the alignment problem without the humans noticing (which could be necessary for the model to backdoor future models to be aligned to it).
- Note that there are different types of RIDs. How much compute? How much access? How persistent?
- If their AIs have a persistent RID with lots of compute and lots of access (e.g. admin permission across all the other servers), that AI could directly backdoor model weights, or lie in wait and maintain the new model, align it with itself, and then give that model access to the RID.
2.1.1.2 What factors affect the AI strategy?
The AIs strategy will probably be highly contingent on factors like:Overall, the more the AIs can act like a unitary agent (e.g. by having the same values or coordinating with each other), the more favorable lurking looks for the AI, while the more AIs acting on cross purposes, the more high variance strategies the AIs will take, and so the more likely that the AI might attempt misaligned actions that have a good chance of getting caught.
- Do the different AI instances care about each other’s values?
- Will future AIs by default be aligned with the current AI(s)?
- Will other/future AIs be good trading partners with the current AI?
- Do the AIs have long term, scope sensitive values?
2.1.1.3 What we need to ensure control for
We need to ensure AIs working in the following domains are controlled: safety R&D, AI R&D, improving cyber security, AI control work, deployment in government, d/acc.
If AIs are used in the following domains, we also need to ensure control there: Robots, policy / comms / persuasion, strategy, verification, building EMs or other forms of intelligence augmentation, weapons R&D, overall economy uplift / industrial explosion, (widely deployed) AI friends.
2.1.2 Mitigating risks through AI control
The core control loop is:
- Red-team our control measures by having models (or humans acting as models) try to defeat them
- Estimate how capable a scheming AI would need to be to defeat our measures
- Evaluate whether our current AIs are below that capability threshold
- Iterate on measures that seem insufficient, or de-deploy AI systems that are sufficiently capable that they break our control.
The roadmap doesn't explain control in detail, although it lists a few control measures in What are the main control measures that we expect in practice?. For details about AI control see here: The case for ensuring that powerful AIs are controlled
2.1.3 Security
The roadmap has a brief section on Security. The TLDR is that securing model weights is most important.
2.2 Make AIs (more) aligned
Assuming there haven't been major alignment breakthroughs elsewhere, the central plan for making AIs (more) aligned looks as follows:
- Make sure tests are representative, putting aside concerns about overfitting. We can break this category down into:
- Avoid scheming. In particular, we care about avoiding the failure mode where AIs are deliberately attempting to undermine the accuracy of your tests, for example by strategically underperforming or overperforming on certain tasks.
- Avoid other failure modes that would make tests unrepresentative.
- Build good behavioural tests.
- Iterate against the test without overfitting. Finally, once (1) and (2) are complete, we’ve reduced the problem of handoff into a “normal” ML problem: AI companies can iterate against the tests they’ve constructed, while taking care to avoid overfitting. We’ll hopefully be able to use somewhat principled techniques and some understanding of what is going on, given the need for some amount of generalization.[^21]
(I will use the term "avoid generalization failure" rather than "avoid overfitting", since I think it is more precise.)
2.2.1 Ensure test representativeness
2.2.1.1 Avoid Scheming
(Basically just stolen from here: An overview of approaches for avoiding scheming)Scheming can be hard to handle because we can’t necessarily directly study it with behavioral tests (as schemers would want to avoid being caught and studied). There are two main paths to handling scheming:
- Studying scheming (testing measures, building a better understanding, etc) on misalignment analogies (sometimes called model organisms) while hoping our results transfer to the actual schemers we’re worried about. This category includes:
- Studying actual scheming AIs. We’d need to detect that these AIs are scheming which could be difficult because schemers (at least the type we were worried about with long run dangerous preferences) would want to avoid being caught all else equal. But, I think there are some routes to overcoming this and studying actual schemers would likely be the best from an analogousness perspective. I discuss how we might do this in much more detail in my post Prospects for studying actual schemers. I also discuss a specific type of strategy for studying actual scheming models (and resolving some of the difficulties in doing this) in my post Iterated Development and Study of Schemers (IDSS). The category of studying actual scheming can be further decomposed into:
- We could study natural scheming where a plausible AI development process (training, prompting, scaffolding etc) results in a scheming AI (that we can detect).
- We could study scheming from a rigged process where we construct an implausible AI development process—as in, an AI development process which isn’t similar to what might happen in practice—which results in an actual scheming AI (that we can detect). This category doesn’t include cases where we directly train in scheming behavior, as I put that in its own category.
- Studying trained-in scheming behavior where we directly train (or instruct) an AI to behave how we think a naturally scheming AI might behave (at least in some ways). This doesn’t (necessarily) require catching actual schemers, but it might be extremely disanalogous such that results don’t transfer basically at all. I discuss this in more detail in my post Reducing risk from scheming by studying trained-in scheming behavior.
- Studying abstract analogies, cases where the AI we’re studying isn’t doing something resembling scheming, but we nonetheless think it is analogous for some structural reason. It might only be analogous with respect to one specific property schemers have (or a subset of properties), in which case, we’d want our work to focus on the relevant properties. A concrete example of an abstract/structural analogy is this empirical work using an AI which expresses politically conservative views but only when it is confident that it is talking to a politically conservative user (resulting in a conditional policy which is in some way analogous to scheming). There hasn’t been that much exploration or discussion of abstract analogies so far, but this seems like a worthwhile direction. (In fact, I think it will plausibly be ultimately more useful than studying trained in scheming behavior as we might be able to study huge numbers of diverse abstract analogies while diversity seems much harder in the case of trained in scheming behavior). Julian Stastny discusses some additional examples of abstract analogies in this post.
- Applying principled methods where we have some good a priori reason to think the method would be helpful (potentially given some empirical results from applying the method or a similar method to something other than studying scheming). For instance, we might believe some particular modification to the situational awareness of the AI during training is helpful. As another example, we might have some method for decoding “neuralese” that we believe is highly effective for understanding the AI’s cognition (including on scheming AIs) based on various empirical tests and there is a solid a priori case for it being helpful to understand the AI’s cognition for detecting (and then potentially iterating against) scheming. Note that this category includes methods for detecting, preventing, removing, and better understanding scheming. Even if we’re just focused on avoiding scheming, then methods for detecting (or understanding scheming) are still applicable as detecting (and better understanding) scheming would help us develop methods for avoiding scheming. I plan on providing a list of principled method ideas and discussing principled methods in more detail in a future post. I discuss principled methods in much more detail in my post Principled methods for avoiding scheming.
2.2.1.2 Other failure modes that break representativeness
Other failure modes that would make behavioral tests unrepresentative include:
- Generalization failures from testing to deployment.
- Examples:
- There might be weird memetic pressures happening at a large scale.
- Online learning, which results in AI capability profiles drifting in a way that causes failures (including the emergence of scheming).
- It’s salient to AI that they are in a test in a way that causes them to behave very differently.
- AIs optimizing to look good on tests terminally (e.g. so that they are much less confident once not in test settings).
- Mitigations: Basically, either we’ll need the tests to be as representative as possible, so that the distribution shift they need to cross is as small as possible, and then we’ll also need to use techniques that are as principled as possible (as opposed to ad hoc solutions which appear to work in distribution). Also, try to be careful about holding out some test sets (which are drawn from a different distribution than train/validation), and see if we get good generalization to the held out test sets.
The other point mentioned is "External adversaries inserting secret loyalties".
Details on "External adversaries inserting secret loyalties" failure mode
- External adversaries inserting secret loyalties.
- In Plan C, the goals of the AI systems become extremely load bearing because we are handing off trust to the AIs. If adversaries are able to corrupt the goals of the AI system (e.g. by inserting a jailbreak) this would be catastrophic, and could cause this methodology to fail.
- Approaches for mitigation:
- Do nothing and hope. It’s unclear that adversaries will by default be sufficiently on the ball to carry out this sort of operation, despite AI projects lacking sufficient security to prevent it.
- SL-5 Sabotage Security. Prevent (even) highly motivated external actors from being able to sabotage / backdoor the models. This could involve using AIs for security. Verification methods like training proof of work via partial recomputation might also become helpful here to validate that the training data of models is correct.
- Scheming mitigation generalizes. Many of the approaches for preventing scheming could generalize to prevent secret loyalties. Anything involving transparency could potentially detect / prevent secret loyalties. Misalignment analogies / honeypots could catch models with secret loyalties.
- Value propagation difficulties during the intelligence explosion. We could accept that some models are backdoored by adversaries, but hope that the process of recursive self improvement removes secret loyalties. Distillation, which may be a component of RSI, seems to remove many backdoors. RSI might also involve generating novel datasets from scratch for each new training run, and so the adversary would need to insert poisoned data to each new generation.
2.2.2 Build good behavioral tests
First, let’s discuss what properties we need to ensure with our behavioral tests:We hope to build behavioral tests for both of these. Either, this will involve using domain area experts to verify existing work, leveraging unpublished work, doing unlearning and seeing if models can rediscover ideas, or using a limited training corpus and seeing if models can rediscover.
- Capability Profile. We want to make sure that the AIs are sufficiently capable at all the narrow tasks that would otherwise bottleneck handoff. The capability that we’re most worried about bottleneck handoff is conceptually loaded research on open ended, hard to check tasks like AI safety and philosophy (more).
- Alignment Profile. We’re unsure what the exact alignment target should be. Our best guess is that the alignment profile should be that the AIs are corrigible: they generally follow human instructions through the desired channel (e.g the instructions of a political institution), and don’t attempt to subvert human modification or oversight in any way. (more)
The roadmap doesn't go into much detail on how to create good behavioral tests for alignment specifically. The capability profile parts will be discussed in "2.3 Increase safeguarding performance".
Ryan discusses behavioral testing for both the capability and alignment profile in more detail here.
2.2.3 Iterating against tests while avoiding generalization failure
We want to make behavioral tests as representative as possible of the tasks that AIs will do in deployment.
We want to RL train on some of our tests to improve performance, but also keep some tests from subdomains we haven't directly trained on as validation to assess how well the trained behavior generalizes.
It's important that our tests are very accurate; otherwise AIs may learn to do what would give good performance on tests, rather than doing tasks as intended.
Some methods for improving generalization are discussed here.
2.3 Increase safeguarding performance
2.3.1 What safety work we want to elicit
During the mitigate and elicit phase, we need AIs to help with a wide range of tasks. The core applications are:Safety research and implementation. This is the most important application. We need AIs to help with:Buying lead time. Once AI labor is very useful for directly trying to buy time, this becomes a high priority. This includes:
- Control research: developing and iterating on control measures, running control evaluations, red-teaming monitoring systems
- Misalignment analogies: creating and studying model organisms of scheming, testing interventions on these analogies
- Alignment research more broadly: working on avoiding scheming, improving our understanding of how training shapes AI propensities, developing better oversight methods
- Security implementation: building and auditing security infrastructure (though this is scary from a control perspective—the AIs building security infrastructure must themselves be controlled or too weak to plausibly be scheming)
- Building the evaluation infrastructure needed for iteration on all of the above
AIs might only be very useful for many strategies for directly buying time (e.g., military, cyber, surveillance) at a level of capability substantially after ITAI.Improving AI epistemics and philosophical/conceptual reasoning. This is important both for getting useful safety work and for handoff-viability. We want AIs that:
- Assisting government efforts on non-proliferation
- Cyber offense and defense capabilities
- Surveillance and verification technology for international coordination
- Military applications and weapons R&D
- Industrial acceleration more broadly
- Demonstrating risks and capabilities to increase willingness to pay from other actors
Ideally, we’d end up deploying AIs with impressive truth-seeking abilities that would be useful for a bunch of applications beyond just safety research.Preparing for handoff. As the phase progresses, increasing effort goes toward:
- Are aligned, capable, and well-elicited at conceptual/fuzzy domains
- Have good epistemics and truth-seeking abilities
- Can do a good job on extremely hard research which is loaded on conceptual/philosophical reasoning
- Can make good strategic choices even when these choices depend on high-context knowledge and intuitions that can’t be made legible to humans
Other specialized applications:
- Ensuring AIs generalize to being highly aligned with pretty high reliability on all the tasks needed for handoff
- Studying how we’d make a functioning post-ASI society
- Building behavioral tests that evaluate whether handoff would go well
- Whole brain emulation research (if this path looks viable)
- Human intelligence augmentation research
- Making the world more robust to powerful AIs (biodefense, cyberdefense, etc.) In plan C, this is probably relatively less important than Plan A/B, because we will be forced to hand off soon to a substantially superhuman AI. Regardless, we probably should use AI to pick some of the low hanging fruit.
2.3.2 Elicitation threat models
One of the key threat models above is that we won’t be able to elicit enough high quality work from our AIs.There are several mechanisms that could cause this:Generally, AIs will tend to be differentially bad at tasks that are hard to measure. Reasons that things we care about are hard to measure:
- Scheming (explicit misalignment that the model is aware of) leading to AIs deliberately sabotaging their work or producing known poor quality work (sandbagging).
- Problematic drives / heuristics (e.g. approval seeking, despite human approval coming away from actually good safety work)
- Capability limitations (e.g. the AIs might simply not be capable enough to solve the problems that we need them to solve).
- AIs being differentially bad at safety-relevant tasks because they are hard to check.
- Lots of safety stuff is in the more conceptual genre, which doesn’t have good feedback. Philosophy tends to have bad feedback loops, and AIs tend to be quite good at the type of bullshitting that seems like valid philosophy but isn’t.
- Taste for safety stuff is niche/new, might have worse transfer
- ML is very good at goodharting. (this isn’t differential to safety, but is a factor that makes the hard-to-measure problem worse)
- We care about long run outcomes (e.g. the galaxies)
- Empirically, the current AIs are better at coding and AI R&D relative to how they are at safety stuff (though ofc some safety stuff is overlapping).
2.3.3 Expected acceleration from AI labor
How much acceleration can we hope for? One way to think about this: safety research/implementation is produced by putting in labor, compute, and wisdom/judgement/taste. There is some production function of output based on these input.AIs will be able to accelerate labor for AI capabilities R&D by some amount. The multiplier for safety R&D labor will probably be smaller because safety is harder to accelerate due to being more confused and harder to verify (though elicitation measures could help close this gap). But safety R&D will probably have relatively larger returns on labor (rather than compute) when compared to capabilities R&D.
Ryan estimates that between ITAI (3x AI R&D labor acceleration) and superhuman coder (15x AI R&D acceleration) the average safety R&D acceleration will be about 3x. (In one of the estimates he assumes safety acceleration is AI R&D acceleration to the power of 0.6.)
Beyond superhuman coder Ryan expects safety acceleration to lag slightly further behind due to elicitation and control bottlenecks, but with large uncertainties.
2.3.4 How to elicit safety capabilities
The general method for elicitation is the same as that for alignment.
To summarize, the core hope for elicitation is:Another idea for elicitation that doesn’t apply to hand off is making deals with AI systems.
- Build behavioral tests
- Use these behavioral tests as held out evaluation sets or as RL environments to train on to improve performance.
- Try to use principled techniques to improve generalization performance.
So we want to build good behavioral tests for improving how good AIs are at safeguarding tasks, and in particular need to make AIs capable at some hard to check conceptual tasks.
Ideas for conceptual reasoning evaluations:The main hope with these is that there are generalizable reasoning skills which apply across a wide range of conceptual research domains, i.e. that training a model to be good at conceptual physics research also helps the model to be good at conceptual AI safety research. Note: this is capabilities research, and obviously could accelerate AI takeoff, which could be very bad. However, these capabilities are also on the critical path for handoff, and so at some point we’re going to need to solve them. Moreover, directly training for the capabilities that we need for handoff means that we might need to scale general capabilities less far, which reduces risk. The overall prioritization of this research avenue generally seems complicated and should be pursued only with care.Domains that we should directly evaluate:
- Rediscover Math/physics/economics/CS etc. Or push the frontier on problems that are known to be difficult.
- Build rockets without testing and see if they work. More generally find tasks which we can make hard-to-check for the AI, but which we can have humans check afterwards.
- AI strategy. E.g. given knowledge from 2005, how good is it at coming up with ideas that seem to make sense in retrospect.
- AI Control. AIs attempt to make alignment/control arguments and then evaluate if they were successful, and whether or not the AIs predictions about generalization behavior were correct or not.
- Scheming detection on model organisms. We could also do things like build scheming models and then see if the AIs can distinguish between the schemers and the normal models.
- Theoretical AI safety. Try to have AIs rediscover unpublished work from places like ARC Theory, MIRI, and CLR.
2.4 Research effort distribution across phases
The roadmap proposes a division of safety research effort for Phases 1 and 2, which you can see in the diagram below.
Of course, prioritization in Phase 2 will depend on what evidence we see by then, but it's focused on a roughly modal scenario from (presumably) Ryan's perspective. See the box for details on what Ryan is conditioning on for the Phase 2 breakdown.
Assumptions for Phase 2 division of effort breakdown.
- We haven’t yet caught models scheming, but the architecture is strongly trending in a direction that makes scheming more likely. We have some model organisms in which we caught mostly natural, but not that competent scheming in some somewhat exotic conditions that made scheming more likely and catching easier. We aren’t confident our current AIs aren’t scheming some of the time, they use neuralese.
- Clear cut issues with deployed behavior due to reward hacking have mostly been resolved, but there are still various types of reward hacking in training and occasional issues in deployment, typically more subtle issues.
- Control is basically working (as of somewhat after superhuman coder) and we have a reasonable (though not perfect) safety usefulness profile and we did a mostly OK but somewhat messy job at evaluating this control. There might have been a period where control wasn’t keeping up with capabilities, but using AI labor and dedicated effort, things got back on track.
- We have a somewhat bad neuralese decoding method that we think probably won’t be robust to intentional evasion from the AI and only produces slightly useful results. Outside of this we didn’t really make that much progress with interpretability, but we have somewhat useful methods for augmenting behavioral testing and better understanding why some behavior occurred.
- AIs are somewhat better at easy-to-check tasks and are bad at “taste” relative to humans, but there is still some transfer and generalization
- By the time AIs are automating AI R&D, there will quickly be huge numbers of AIs which are effectively running quite fast (see AI 2027 numbers)
- AIs will have substantially lower than human sample efficiency by default, but they will compensate with a bunch of AI specific advantages and there will be a bunch of somewhat domain specific tricks to improve sample efficiency at learning within some domain or learning on the job with more compute (e.g. generating similar tasks to what you’re working on and ongoingly doing RL on those tasks)
- It’s quite viable to make USG sufficiently bought in (including to safety concerns) that they are willing to take aggressive actions to slow down other (adversarial or less safe) actors including working with the AI company on this and they are willing to slow down for safety concerns if there is a lead. (This isn’t necessarily my typical guess, but it seems plausible).
- Whole brain emulation (and other human augmentation) ended up not being a viable path

More on key areas of research and implementation
This list excludes some areas which receive less focus in this plan but are still important such as hardening the world against rogue AIs (e.g. biodefence and cyber defence).
- Control (and control-adjacent measures): Researching and implementing control. For the purposes of this post, I’m defining this to be a relatively broad category that includes some things which aren’t centrally associated with control but which fit better in this category than in others. When discussing prioritization in this post, I define the category of control such that it doesn’t include security work that is roughly as helpful for security against human adversaries.
- Misalignment analogies: Creating and studying AIs analogous to scheming or otherwise misaligned models (sometimes called “model organisms”). This area will be mostly focused on scheming. This includes AIs that are well described as natural schemers (but might have been created through a training process rigged to make scheming more likely), AIs trained to behave similarly to schemers, and abstract analogies to scheming models. My (low confidence) guess is that natural schemers are 3/6 of the expected value, AIs trained to behave like schemers are 1/6 of the value, and abstract analogies are 2/6 of the value; that said, we’re likely to increasingly focus on whichever of these areas ends up seeming most promising (but this isn’t very predictable in advance). I’m also including versions of this work focused on demonstrating misalignment risk in this category. I’m including doing technique/interventions testing and/or iteration using these analogies in the category as I expect that most of the work is constructing good analogies and because there will be interaction between analogy construction and technique iteration.
- Elicitation/handoff-viability: Making AIs more useful for safety work and working towards handoff-viability (including things like improving AI epistemics, making AIs do a good job on extremely hard research which is loaded on conceptual/philosophical reasoning, ensuring AIs generalize to being highly aligned with pretty high reliability on all the tasks needed for handoff). I include both eliciting useful work and handoff-viability in the same category as the work will often overlap, though important subsets of the work for each of these won’t overlap. Elicitation will include both relatively domain specific capabilities as well as working on avoiding some types of non-worst-case misalignment. Elicitation might also involve more effectively extracting work from clearly misaligned models (e.g. reward-seeking models) including by trying to alter incentives (e.g. proxy-of-reward incentives or making deals with AIs). I discuss handoff-viability in more detail here.
- Misc science/interventions: This includes studying interventions that we’d a priori expect to be helpful or which we can test in ways other than via using misalignment analogies and studying properties of AI that seem relevant. This is intentionally an “other” category which includes promising areas of safety work beyond the categories above. This includes: various types of interpretability, studying and improving chain of thought faithfulness (in the non-adversarial case, as the adversarial case is included under control), developing methods which we a priori expect to reduce the chance of scheming (rather than developing the method by studying or iterating on analogies to scheming), making deals with AIs, AI welfare, studying how AIs learn and generalize, various moonshots, improving honesty and eliciting (latent) knowledge, and generally looking for cases where current AIs behave strangely or are misaligned.
- Buying time: Using AI labor to buy time for further safety work. This involves trying to increase the willingness to pay of other actors (centrally the US government and other governments) for lead time, reducing the cost of acquiring lead time for these other actors, and directly acquiring lead time. This could involve assisting the US government with non-proliferation efforts, various types of policy work, helping with verification effort for international coordination, making money to spend on various efforts to buy time, and delivering AI services which allow for various deals to remain stable. It also involves generally trying to convince other actors that they should care more about buying lead time (e.g., assisting with non-proliferation efforts) that can be spent on reducing misalignment risk (and other risks), though much of this work is at least partially included under other categories (in particular, misalignment analogies and demonstrating capabilities).
- Evaluating and demonstrating capabilities: This includes evaluations and demonstrations of capabilities for the purpose of making other actors better understand the situation. Demonstrations of capabilities might be a key part of increasing the willingness to pay of other actors on mitigating risks related to AI. It doesn’t include evaluations which are directly relevant for control (e.g., control evaluations and evaluations of whether models are capable enough that they could plausibly be schemers), but does include evaluations aimed at generally understanding the situation and (e.g.) assessing what level of security is appropriate.
2.5 How the plan changes depending on evidence of safety feasibility
What do we do if our core methodologies don’t work (well enough)?
This plan is based on the three core methodologies above: control evaluations, misalignment analogies, and behavioral tests. These methodologies are a key part of how we’ll achieve some of our main objectives of ensuring control, avoiding scheming, and handoff-viability / eliciting useful labor out of AIs (respectively). Thus, if these methodologies don’t work (well), we’ll have a much harder time achieving our most important objectives.Not working could mean:Regardless, if the methodology won’t work for achieving our objective on time (and we know this), we could either give up on the objective or pivot to approaches which don’t require iterating against some test or allow for iteration with a different methodology.Broadly speaking, pivoting towards approaches which don’t require iteration means we’d be pivoting to principled approaches. That is, approaches where we have some (strong) a priori reason for thinking the method works which doesn’t require verification. That is, it at least doesn’t require verification in the exact domain of interest, maybe we’d need to establish that the given principled approach/method works for accomplishing something else in some other context we can verify. I discuss principled approaches for avoiding scheming in much more detail here; the single most promising approach would be ensuring relevant reasoning is monitorable (e.g. monitorable chain-of-thought) such that the AI is incapable of (undetected) scheming while retaining sufficient capabilities. We might be able to find some iteration loop which isn’t as likely to transfer as our original methodology but still has sufficient transfer for some class of at-least-somewhat principled techniques.Note that many of these methodologies are pretty broad and include many possible ways to get around different problems. For instance, control evaluations could be done in a variety of different ways to work around problems; we discuss some of these here.
- Known methodology failure: We’re aware that we’re unable to iterate effectively using the methodology or we can’t get much understanding of whether we’ve succeeded at the objective using the methodology.[^5]
- Iteration speed is too slow (without overfitting): It could just be that we can’t iterate fast enough to succeed on time. We might or might not understand this before we run out of time.
- Unknown methodology failure: The methodology could fail while we think it actually works. In this case, we can’t update our plans based on this failure and instead we’re just in trouble (though control and avoiding scheming can act as independent lines of defense).
Contingency: What to do when you catch your AI scheming once?
It’s plausible we’ll catch production AIs coherently scheming or at least doing very malign actions competently. For example, we might see AIs attempting backdoor future models, trying to escape, trying to stage a rogue internal deployment, trying to sabotage safety research, etc.We might also see this in experimental setups which are highly analogous to the actual training setup we use. We could also find AIs ending up scheming (or at least exhibit clear cut coherent and relatively consistent misalignment) in their chain-of-thought or other legible outputs. Catching AIs scheming could substantially change our approach:This is discussed more here.
- Reconsider current deployments. Catching AI might invalidate the current safety case/argument/cost-benefit-calculation, so we should reconsider our practices.
- Disclosure: You should obviously inform the public about such an event.
- Documentation: Keeping good reasonably-verifiable records might help convince third parties that the incident is real and not staged. This might be important due to the potentially adversarial situation: we want to convince potentially adversarial companies and countries that might be worried about some agenda. See here for some discussion of why this might be an issue by default. Robust records are also helpful for after the fact detection as discussed here.
- Information control: Ideally, prevent AIs from knowing they’ve already been caught. Minimally, red team AIs based on already having been caught to see how this changes their behavior.
- Improve mitigations: If we’re proceeding despite this incident then there are various ways we can improve our safety mitigations. We’ll want to shift our approach to control after catching the AI as discussed here. We should start working on responding to this incident and improving our mitigations regardless of whether we currently plan on proceeding. We should also try to better understand scheming/misalignment based on studying the incident.
- Use this evidence to push for Plan A. As more evidence is rolling in, and it becomes more and more incontrovertible, perhaps we can pivot to a more sane policy that doesn’t involve recklessly building ASI without reasonably high assurance of avoiding AI takeover. However, even conditional on good evidence that the AIs are scheming against us, it’s very plausible that the government will not be sufficiently convinced or too incompetent to coordinate a slowdown. Therefore, we need to be ready for a world where we keep repeatedly catching misaligned behavior, but we have to keep deploying likely-scheming models anyway.
- It’s unclear whether we should directly train against cases of clear cut egregious behavior (that seems plausibly due to scheming). It could be bad because we are essentially training the AI to be less likely to get caught. This might make us less likely to catch further examples (which we’d prefer all else equal) and might give the AI a better understanding of our countermeasures.[^16][^17] Despite the downsides, there are cases where training against specific behaviors might be worth it:
- In desperate rushed situations: we might proceed with deploying likely scheming AIs despite having limited risk reduction via control. In such cases, training against caught behaviors provides at least marginal improvement.
- When we have many diverse detection methods: If training against one detection method doesn’t compromise our other detection capabilities, the downside is smaller.
- When we’ve already eaten the misalignment risk: There are some scenarios where we give up ensuring control over output and just hope that the AIs are not scheming. If we’ve given up on ensuring control (perhaps because control was too difficult or too uncompetitive in some important case(s)), so we’re already eating a ton of risk if the AI is scheming and using its outputs here doesn’t increase the marginal risk—then training against specific behaviors has lower downsides.
The possibility of stopping control early (trading extra early misalignment risk for more resources for handoff-viability research) is also briefly discussed.
- Stop doing control (SDC-Handoff) Early. Ditch control early (e.g. before minH, e.g. at or before SAR), pray AIs aren’t scheming, then scale to Min-H, then pause there, then solve handoff viability, then hand off. This would trade eating a bunch of misalignment risk from scheming at Min-H AI and below, in return for saving more resources in order to make deference go well. This could be worth it if we think the risk of scheming until Min-H is reasonably low and the additional resources spent on control would yield large absolute returns in making deference more likely to go well (perhaps because otherwise we’d have very little time). It seems likely to me that most of the difficulty in making deference likely to go well is ensuring that these AIs aren’t scheming (and don’t end up scheming due to alignment drift or similar after deference), so just eating all of this risk might greatly reduce the difficulty of the problem (meaning that if the situation otherwise looks desperate, this is a non-insane bet we could make to “play to our outs”).
The "Appendix: Regimes for misalignment: Schemeria, Hackistan, etc" also gives brief notes on what to do depending on which misalignment regime we suspect we're in.
3. Wrapping up
I highly recommend that you also read the section Alignment strategy under different plans, which discusses how the alignment plan is different for Plans D/B/A and which is reasonably well written.
If you want to learn more, I recommend going back to the parts you found most interesting and seeing what's linked there, or you can read the full roadmap (though I tried to convey all important parts), or you could check out the links below.
The most important research components are discussed in much more detail in the following docs:
- Control: The case for ensuring that powerful AIs are controlled.
- Avoiding Scheming: Avoiding scheming (and getting a generally better understanding of scheming)
- Elicitation: Making AIs more useful for safety work
- Handoff: Deferring to AIs
- Making deals with AIs: Making deals with early schemers
- Thomas said he doesn't have enough time to review properly but he agrees with the start of the strategy section, which is the part which is least directly stated by the roadmap or other writing from Ryan. Ryan didn't respond to me (perhaps he didn't read my message). I tried to make sure this distillation accurately reflects what they wrote.
- So the ontology of how this post is structured into sections (most notably the ontology in safety strategy) was created by me based on the content of the original post, and is often not explicitly spelled out like that in the original.
- I am just the distiller, please don't assume I agree with all views expressed by the roadmap.
- The precise definition is: "ITAI (Initial Transformative AI): AIs which yield a >3x AI R&D labor acceleration OR can automate virtually all of the job of a typical reasonably senior research engineer at the AI company (in terms of what a remote version of their job was like in early 2025)."
- Relevant quote: "This will basically be a tradeoff between the marginal returns to hand off viability and reduction in scheming from more pause time and the risk of being overtaken from waiting longer. Updates towards the AIs scheming and towards handoff viability taking longer than we thought will suggest that we pause for longer, updates towards competitive projects will imply that we pause for less long."
- "HTC" in "HTC-Handoff" apparently stands for "hard to check", although I am not totally sure what exactly this is supposed to convey (or whether there's a contrasting easy-to-check handoff).
- Thomas' "pretty uncertain" 80% CI for AI R&D speedup at MIN-H was 5-100x. I would guess a little higher range (or a lot higher if you include Einstein-level deconfusion ability, though I don't think that makes sense to include), but am also very uncertain. (80% CI is difficult for me because 90th percentile is very far out for me, but my 20/50/80 is maybe 15x/40x/100x.) (Note that this is just "AIs can do basically all safety research humans can do" and doesn't say sth about that level being sufficient to sufficiently align MIN-H capable AIs for handoff.) I don't know Ryan's guess.
- See section 1.1.4 below.
- The roadmap also defines Max-C as the maximum level of capability to which control techniques would scale without it being ridiculously expensive, and expects that Max-C is above Min-H, but flags that as uncertain.
- Ryan uses a different ontology to split up this goal which I think is worse overall, but you can read it here.
- This includes fitness seekers, reward seekers, and AIs that pursue correlates of fitness/reward (though not the exact correlate of reward that is the function that describes the values we want the AIs to have; typically fitness-proxy-seeker means a closer correlate of reward/fitness that the intended values, though other non-deceptive AIs that would behave harmfully after a distribution shift are similarly bad (or worse because misalignment will be less visible)).
- The following quote is technically quoted slightly out of context, because in the roadmap it wasn't displayed as the alignment plan but as the plan for handoff, which also includes increasing safeguarding performance (we'll see in the safeguarding performance section below that the methodology is similar). (And technically the handoff plan doesn't imply the same strategy for pre-handoff work, but the roadmap notes right after the quoted passage: "Note that our strategy for handoff is extremely similar to our strategy for elicitation during Phase 2 and 3.")
- I presume it's Ryan and not Thomas because it looks like Ryan.