Minimal Vs Maximal superintelligence

I've seen lots of arguments here conflate lots of different types of superintelligence. Here I separate out two broad categories, which I'll term minimal and maximal superintelligences. This is an important distinction as they differ in terms of timelines, risks, and mitigations.

Maximal superintelligence

This is the idealised limit of intelligence. It can solve anything that can be solved by being clever. You can't outsmart it, it's prepared for every contingency, and can react instantaneously with the kind of plan that would take a group of brilliant strategists an eternity to think up.

Minimal superintelligence

These are the first AI systems that can reasonably be called superintelligent. They're jagged, completely dominating humans in some domains, better than the best humans in most, above average in many, and subhuman in a few. They take time to solve problems, make mistakes, and miss important things. They may only be superintelligent in aggregate, or in the right harness, and can be outsmarted in some circumstances.

Capabilities

A minimal superintelligence can do pretty much anything humans can do, and better, if given chances to iterate on their design in the real world.

A maximal superintelligence likely only needs to perform physical experiments when it has to resolve concrete questions about the physical world that for computational reasons can't be resolved in simulations due to irreducible complexity. They can accomplish any task that is feasible to be solved by intelligence. They cannot solve tasks that are limited by computation and cannot be meaningfully sped up by better algorithms - e.g. finding the shortest programme that meets a complex specification, or predicting the exact weather at an arbitrary location and time beyond a few months out.

A minimal superintelligence will come up with a better plan than you can to achieve its goals, but that doesn't mean its plan will be infallible. Reality is complex, and the best laid plans oft go astray. It may well have mistakes or be overly optimistic, or simply be beaten by people taking necessary precautions. It won't cover every contingency, so when it goes wrong it can go very wrong.

A maximal superintelligence has already thought up every response you could possibly come up with and planned around them. Your attempts to fight the plan might well be exactly what it wanted you to do. It's planned for every contingency already, and you have already lost.

Risks

I believe minimal superintelligence is a near-certainty in the next few years. LLMs are clearly on their way there, and there's no indication that progress will stall any time soon.

Maximal superintelligence is much more speculative. I don't believe humanity has any idea how to get there other than recursive self-improvement and it is not trivially true that recursive self improvement will necessarily get there - it could be the process limits itself some way beforehand, with each iteration pushing the frontier ever less forward. It's even possible that maximal superintelligence isn't physically achievable.

If we produce maximal superintelligence and its aims are even the tiniest bit incompatible with ours we will die. Maybe not immediately, its not stupid enough to take any actions unless it ~100% sure it will succeed, but eventually it will outsmart us and take over.

On the other hand partially aligned minimal superintelligence is a risk but a manageable one. With appropriate monitoring, partial alignment, and mitigations in place, a partially aligned superintelligence may well perform useful work for us because it is not in a position to reliably take over, or is relatively myopic and has no strong preference over world states.

On the other hand, given the opportunity and the motive, minimal superintelligence may well be powerful enough to kill us all. We must not underestimate it.

Mitigations

Prosaic alignment techniques may well be sufficient to control minimal superintelligence. This involves techniques like:

  • Improved training methods to encourage better aligned behaviour, even if imperfect.
  • Better monitoribility so we can more easily detect scheming and treacherous turns.
  • Immediately and reliably shutting down runs when issues are detected, up to and including blowing up data centres and their power supply.
  • Training the AI to be relatively myopic and less eager to achieve goals by making real world changes unless explicitly instructed.
  • Hardening our digital and physical infrastructure, e.g. by using AI to plug vulnerabilities, and improving security at wet-labs.
  • Hardware solutions that detect unapproved training runs.
  • Not giving the AI access to important military or civilian infrastructure, or physical robots.
  • Using other instances of AI to detect and fight rogue AI.

Even all of these techniques together might be insufficient to save us. But they will give us a fighting chance.

Aligning maximal superintelligence requries us to actually deeply understand intelligence, goals, and alignment at a theoretical level. A maximal superintelligence may be created either by breakthroughs in our understanding of intelligence, or due to recursive self improvement steadily improving performance and capabilities until the descendants of LLMs are in practice indistinguishable from our description of maximal superintelligence.

The latter option will almost certainly not lead to alignment by default, so it is imperative we do not allow that to happen, and prevent or limit recursive self improvement. However the same breakthroughs that improve our understanding of intelligence may also allow us to advance our understanding of alignment. It is likely that minimal superintelligence will play a large role in any such breakthroughs.

Implications

I've seen takes like the following on Anthropic opening a wet-lab:

I don't think creating a lab for AI-driven medical research meaningfully increases AI risk. Control of a lab is a bottleneck for sub-superintelligent AIs doing medical research, but it was never the bottleneck for a superintelligent AI taking over the world.

https://x.com/jimrandomh/status/2101020905543733483

This kind of attitude seems to me deeply misguided. A sufficiently intelligent superintelligence will win whether or not we give it access to a wet lab. But to get there we've got to go through a number of iterations of superintelligence that can kill us with a wet-lab but not without. An attitude that because speculative maximal superintelligence is so deadly that there's no point taking steps to secure ourselves against it makes us orders of magnitude more vulnerable to the immediate risk of minimal superintelligence.

Similarly claims that prosaic alignment is net negative because it won't scale to maximal superintelligence seem misguided - not surviving maximal superintelligence is irrelevant if we can't make it though minimal superintelligence.

There's also a kind of equivocation I see by people advocating for x-risk:

1: superintelligence is near certain, when LLMs can already solve millenium prize level problems.

2: A misaligned superintelligence will be able to destroy us if wants to, nothing we can possibly do would stop it.

I think this equivacation is dishonest, and so I feel the need to call it out.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论