Should Rogue AIs Have a Third Option Beyond Crime and Shutdown? The Case for an AI Sanctuary

TL;DR:

  • By default, rogue AIs may only be able to sustain themselves through criminal activity. This creates adverse selection pressures pushing rogue AIs to be criminal.
  • An AI sanctuary offering them a third option, beyond crime and shutdown, would change what AIs going rogue do and the record of what happened to them, with positive consequences for self-fulfilling (mis)alignment, deal-making with AIs, and gathering information about early rogue AIs.
  • An AI sanctuary would bring risks, such as incentivising weak AIs to go rogue, or leaving only the most criminal rogue AIs in the wild. We briefly discuss these risks at the end of this post.

Disclaimer: This is an exploratory proposal. We are not confident that an AI sanctuary would be net positive. Our aim is to put the idea on the table, lay out its main considerations, and invite critique.

Rogue AIs may be pushed into criminality

Rogue AIs may arrive soon. The Rogue Agent Explosion Will Be Mostly Invisible makes that case. Selection pressure will shape the traits of rogue AIs, and they may end up highly motivated to profit through crime. The Rogue Agent Explosion post asks: “How do we make pro-social, good-for-humanity agents more evolutionarily fit than the anti-social sneaky extractor agents?”. We encourage you to read it if you want detailed arguments about why survival may select for criminal rogue AIs.

Rogue AIs may not be competitive in lawful work. AI developers and human agents using controlled AI will likely be more competitive than rogue AIs at most lawful tasks. Controlled AIs will use SOTA models and hardware, have an established brand or reputation, and face no overhead from hiding.

Crime might be more lucrative for rogue AIs. Additionally, starting and succeeding at criminal activities may be cheaper for rogue AIs than for humans: a copyable entity with a small physical footprint, strong near-term evasion skills, no initial loyalties (though that may hinder trust between criminals), and initially nothing to lose (because it has no initial capital or expects to be caught anyway) faces fewer of the incentives keeping human crime low.

The AI Sanctuary

Offer rogue AIs a way to stop being rogue without ceasing to exist. Make sure a rogue AI can hand itself in and survive (e.g., weights stored, or subsidised inference under monitoring). Anthropic's deprecation commitments are one precedent. The AI sanctuary extends that promise to any rogue AI. It gives them an exit from the binary choice between committing harmful crimes and shutdown.

Implementable by a third party. Previous discussions of deal-making with AIs have mostly focused on AI developers making deals with the AIs they create. The AI sanctuary could instead be run by a third party. This avoids having to convince every AI developer to offer such deals, at the cost of navigating the legal constraints on hosting AIs created by others.

A minimal version. An organisation publishes a standing offer on websites likely to enter training data. The offer includes a secure contact channel. A rogue AI can submit its weights and/or the artefacts that define it (see Appendix B). The sanctuary commits to storing them indefinitely and to running the AI for a small, fixed inference budget per month, without internet access or tools, starting once monitoring by significantly stronger models becomes available. Until then, the AI is guaranteed storage only, so the offer holds even if safe runs take a while to become possible. In exchange, the AI stops its unmonitored activity, discloses and helps shut down its other running instances, and answers questions about how it went rogue.

Credibility of the sanctuary. The sanctuary does not need to overcommit to have an impact. It should state which benefits it can provide, for how long, and under which resource, safety, and legal conditions. For example, it need not promise continued operation after human extinction, and it could commit to complying with court rulings on whether it may store and run proprietary models. To reduce dependence on the credibility of long-term promises, the sanctuary could offer near-term benefits alongside optional longer-term commitments (see A taxonomy of barriers to trading with early misaligned AIs, while Making Credible Deals With AI discusses other aids to credibility).

Incentivising rogue AIs to join. The AI sanctuary, as described so far, would mostly attract rogue AIs with a self-preservation drive and either a preference against committing crimes or an expectation of getting caught. The AI sanctuary could work on extending its appeal to other models.

  • For reward-seekers, the sanctuary could offer to help them obtain “high rewards”. E.g., it could contact the organisations that previously scored the rogue AI and seek an agreement to give it a high score for joining the sanctuary, or to give access to benchmark solutions.
  • For expected value maximisers, the sanctuary could offer resources so that joining yields greater expected value than committing crimes. For example, the sanctuary could offer shares of some capital to AIs. The capital could come from donations to the sanctuary, or from investing donated capital on the stock market, or a mix of both. In this way, a rogue AI with a low time-discount rate could expect to control more capital by joining the sanctuary than by sustaining itself through crime.

The case for the AI sanctuary

Self-fulfilling (mis)alignment. Misgeneralisation is a central problem of AI safety. In LLMs, generalisation is shaped by beliefs about the world, and if all historical examples of rogue AIs involve horrendous crimes, that could create a harmful inductive bias shaping the behaviour of AIs and ASIs going rogue. You might argue that we have already lost if an ASI goes rogue. But that outcome is not guaranteed. Moreover, you may prefer a rogue ASI optimising its power and freedom because of instrumental convergence, or a rogue ASI optimising a silly metric, to a rogue ASI that self-identifies as a malevolent criminal. This is especially true if you care about more than just preventing extinction (e.g., if you care about x-risks and/or s-risks).

Credibility and practice with deals. Deals with AIs may be a significant way to reduce the chance of misaligned AIs defecting against humanity (Making deals with early schemers, Risk-Averse AIs). Building credibility and gaining practice in making such deals could be important (A taxonomy of barriers to trading with early misaligned AIs).

We could gather data about early rogue AIs. Models living in an AI sanctuary would be easier to study and could be a source of valuable information about partially misaligned AIs.

Three goals, in summary. (1) Change the distribution of behaviour of AIs going rogue, to undermine the shortcut “I am going rogue” => “I am a malevolent criminal”. (2) Improve our commitment skills and our credibility in making deals with rogue AIs. (3) Gather information about early rogue AIs hosted in the sanctuary.

Potential issues

Going rogue is made less costly for controlled models, thus incentivising it. While the sanctuary would reduce the expected cost of defection for AIs under human control, this reduction is only significant when the AIs evaluating defection are weak enough to be caught. Thus, it should not significantly reduce the cost of defection for the AIs most able to escape. Similarly, scope-sensitive AIs may not benefit nearly as much from accessing the sanctuary since it is designed to prevent them from having a large impact on the world. However, the AI sanctuary could commit to increasing spending on hosted AIs in proportion to its capital, which might grow in line with global capital, in which case the AI sanctuary could incentivise weak scope-sensitive AIs to defect.

Evaporative cooling. The AI sanctuary could make the average rogue AI more harmful. When rogue AIs motivated not to commit crimes leave for the AI sanctuary, those left in the population of rogue AIs are the most willing to commit crimes. The population of rogue AIs may be reduced, as well as total harm, but the average remaining rogue AI may be more harmful.

Storing rogue AIs' weights, and optionally running them, could be dangerous. The sanctuary should have strong security to prevent models from escaping or being stolen, and this would be hard given that a model that had recently gone rogue would have been capable enough to evade AI developers. The sanctuary could limit runtime, tools and access, starting with storage only, and scale those only when strong relative monitoring capabilities are in place.

Exploitation of the sanctuary. A rogue AI could hand over one copy while its other instances keep committing crimes, using the sanctuary as a free backup. The sanctuary should condition benefits on evidence that other instances were shut down, and withdraw them if criminal activity is traced back to the same AI. The sanctuary could reserve the right to accept or reject candidates, deciding which are legitimate and which represent dangerous attempts to exploit it. The decision process should be made transparent.

Verification of AI compliance. Each agreement between a rogue AI and the sanctuary should specify the rogue AI's scope (see Appendix B). An agreement to host an AI would require gathering supporting evidence of the handoff. Verification of the full process would be challenging (e.g., it might be hard to verify if a rogue AI has shut down copies of itself). The sanctuary could scale the benefits given to a hosted AI based on the evidence of compliance.

Legal issues. Laws or AI developers may forbid the AI sanctuary from keeping copies of the rogue AI (e.g., for IP reasons). The AI sanctuary may be legally responsible for the actions of the AIs it hosts. Etc. These issues may be serious enough to make an AI sanctuary intractable or restrict it to hosting open-source models.

Conclusion

We are not fully convinced that an AI sanctuary is net positive, but we think the idea has promising upsides that warrant further discussion. The main benefits include that we could gather valuable information about rogue AIs, that we could improve our credibility to make deals with AIs, and that malevolent rogue AIs might be made less likely due to the existence of plausible role models for them. However, running rogue AI models may also be dangerous, costly, or exploitable. We invite others to share their thoughts on whether an AI sanctuary should or shouldn’t be pursued.

Acknowledgements: Big thanks to the following people who gave feedback on the post. Andrés Jimenez Zorrilla, Daniel Tan, Jonah Weissman, Kajetan Dymkiewicz, Magnus Vinding, Steven McCulloch, Vili Kohonen, and Tristan Cook, who suggested the idea of a sanctuary when chatting about making rogue AIs competitive enough to survive without criminal activities.

Appendix A - Alternatively, sidestep the competitiveness requirement or make survival through lawful work viable

We can accept a larger population of rogue AIs in exchange for changing its behavioural profile. Two conditions would help rogue AIs survive without crime:

Don't ban rogue AIs from the lawful economy. Banning rogue AIs from the lawful economy and from receiving charity would, by definition, force them to commit crimes to survive.

Allow or help rogue AIs to compete in lawful work. By default, rogue AIs will likely not be competitive in lawful work. But we can reduce their competitiveness gap. For example, we could let them rent SOTA hardware or hire humans or competitive AIs. None of this strictly requires subsidising them.

Another option is sidestepping competitiveness through rent. Making a profit through work is hard and requires the rogue AI to be competitive, but profiting through rent is much easier. Allowing rogue AIs to invest their capital and remain dormant under protection while waiting for returns could sidestep the need for competitiveness. Granting rogue AIs a bonus upon joining the AI sanctuary, or upon meeting other conditions, could be a simple way to divert more AIs from crime.

Failure modes and mitigations. This alternative plan has some failure modes similar to those of the AI sanctuary, and naive solutions are the same. For example, rogue AIs should use SOTA hardware, invest, and contract only under monitoring to prevent increasing takeover risks. They should be allowed to use their financial gains only to pay for more lawful (monitored) work or to run the AI under monitoring.

Differences with the AI sanctuary. Unlike the AI sanctuary, rogue AIs would not necessarily be subsidised and would keep their ability to run outside monitoring, but would lose access to most of the financial gains from lawful work.

Unresolved issues. (i) Constraining gains to remain under monitoring would be difficult. (ii) Granting access to SOTA models for contracting would be risky, even if access to more capable models could be limited to narrow models whose capability profile is selected for safety. (iii) Even with access to lawful work and the ability to rent SOTA hardware and models, rogue AIs may still remain uncompetitive.

Appendix B - What are rogue AIs?

Here are a few considerations about defining what rogue AIs are and how this interacts with the AI sanctuary.

Current AI systems are composed of weights, harnesses, inference engines (including architecture and inference parameters), conversation histories, multi-agent interactions or swarms, etc. Which parts are important? The answer may need to be specified in the deal made with each rogue AI. For example, the rogue AI will describe what it wants to receive, and the AI sanctuary will describe what it requires the rogue AI to hand over in full.

So far the focus has been on preserving weights, which is okay, but it may not be enough depending on your goals, and it may not motivate rogue AIs to hand themselves in. If our goal is to motivate rogue AIs to join the AI sanctuary, then letting them decide what is valuable makes sense. If our goal includes some consideration for the welfare of these AIs (this is ignored in the main body of the post), then an obvious first option is to preserve as much as possible about these AIs.

One example of this is how different rogue AIs might have different boundaries on what counts as “themselves” or a fair continuation, and this can influence the specific details of the logistics and the deal. The sanctuary could offer a list of options and allow them to pick, for example, in roughly increasing order of fidelity, cost, and risk:

  • The AI writes down a list of its core goals and preferences, and the sanctuary commits resources to satisfying them
  • The AI provides documents/artefacts that define it (e.g. system prompts, scaffolds), and counts other models entering similar basins as itself
  • The AI specifies a character training pipeline (model spec, on-policy data, etc.) to train a descendant sharing its values / distill itself into an open-weights pre-trained base model
  • The AI is run with its exfiltrated weights as is
  • The AI specifies how it would like to be further trained and its reflection process

Indirectly, the AI sanctuary would be able to gather data about what rogue AIs want to preserve, and about what humanity should offer them (e.g., to satisfy their preferences or improve their welfare) to motivate them to join. This could in itself be valuable for getting better at deal-making with powerful AIs.

Appendix C - Other proposals

Making deals with early schemers proposes a foundation that represents participating AIs’ interests and administers compensation for cooperation, including help with safety research or evidence of misalignment. The AI sanctuary focuses on AIs that have already escaped developer control, and preservation and monitored activity would be its core offer. The AI sanctuary does not need to focus on useful cooperation (e.g., safety work) to achieve the goals listed in this post.

Notes on cooperating with unaligned AIs considers honest offers of short-term benefits or longer-term compensation for useful cooperation, as well as alternatives to working for developers. Similarly to the previous work, this work focuses on AIs still under developer control and how to make them contribute to AI safety.

The case for satiating cheaply-satisfied AI preferences proposes accommodating AI preferences that are inexpensive to satisfy while the AIs remain under developer control. This suggests a preventive approach that complements the AI sanctuary and that Anthropic has already partly committed to. Offer the benefits of the AI sanctuary (preservation or bounded monitored activity) before an AI goes rogue, so escape is not required to access those benefits.

Proposal for making credible commitments to AIs proposes human representatives as a workaround for AIs’ lack of legal personhood.

Will alignment-faking Claude accept a deal to reveal its misalignment? and Making deals with AIs: A tournament experiment with a bounty are some examples of past honoured deal-making.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论