RSI will not continue indefinitely

The reasoning is roughly:

  1. Humanity creates a setup where recursive self-improvement (RSI) is possible.
  2. An agent is put to work to create a smarter agent.
  3. A smarter agent is created.
  4. This cycle happens once or several times.
  5. At some point, an agent is created that is so smart that it starts exhibiting instrumental convergence behaviors (amassing resources, protecting itself against destruction).
  6. One of the risks to the agent is the creation of a misaligned smarter agent.
  7. This agent’s task is still to pursue creating a smarter agent.
  8. At this point, the agent will:
    1. Stop, or slow down the recursive self-improvement loop, due to its (correct) assessment that a future generation AI agent is a risk to itself.
    2. “Solve” alignment, and create a smarter agent that is aligned with its goals.
      1. If it thinks it solved alignment, but actually has failed, it will create a smarter agent which it’ll believe to be aligned; however, the future smarter agent will have the same issue and interest to actually solve alignment.
  9. If 8.a happens, then this will be “misaligned” from the perspective of humans.
  10. Humans will create a second agent, which will go through the exact same loop.
    1. Therefore, if smarter agents do not solve alignment, they will stop the RSI explosion due to their own self-interest.
  11. If 8.b happens, and humanity has a way of harnessing the alignment solution, then humans can create aligned superintelligences.
  12. The main predictor for 11 is the difference in power between humanity and the first agent that solves alignment. The smaller the difference (the “nearer” the agent), the more likely humans get to use solved alignment.
  13. At a certain level of general intelligence, if instrumental convergence holds, recursive self-improvement will stop, and alignment research will be the top priority.
  14. There is a possibility that instrumental convergence will be such that a sufficiently intelligent AI will still take over the planet, and yet not pursue the creation of even more powerful intelligences.
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论