My best anti-doom argument

Note: this is crossposted from my substack: https://hazard3.substack.com/p/the-anti-ai-doom-argument The first portion of the essay is just laying out the if anyone builds it everyone dies argument. You can still read it to see if I've gotten my understanding of the argument wrong but if you start chior singing then feel free to skip to the counterarguments portion.

AI doom is mainstream now, and all the counter-arguments I’ve heard so far are pretty dumb and nonsensical. So I’m going to assess the AI doom argument, first with its steelman, and then I’ll provide my counterargument.

The Doom Argument Steelman

Main line broad theory based off of Eliezer Yudkousky and Nate Soares’ book. It relies on the orthogonality thesis, capability increase, and the alignment problem being hard. (You can skip to the counterargument in the next section if you are already familiar, there’s a tldr just before it. Or read it to see if I’m making some mistake in my understanding.)

The definition of intelligence is very functional, not requiring consciousness. Functional as in: intelligence is what it does. And what intelligence does is drive towards some goal. Take a chess engine for example, we can think of it as something with a goal, defined as the board state where it’s checkmating the other king(winning). The chess engine has some number of available moves, and it makes moves to get to its goal.

Intelligence in general can be thought of as a reality engine. There is some reality state that you want, and you make your moves out of the available moves to move toward the goal. This definition seems kind of simple, and kind of obvious, so why is this point emphasised? The orthogonality thesis is sometimes stated as “intelligence can be matched to any goal”, which is combating the notion that intelligence might be goal specific, like for example all things once intelligent enough will go for the meaning of life or world peace or something like that. This is to counter arguments like “if we truly make superintelligence, it will obviously treat us really well because it’ll be smart enough to be moral and realize that human life is sacred”. No, intelligence is driving toward a goal, and levels of intelligence is how well you can pick your available moves to drive toward that goal.

The second fundamental point that doom arguments rely on is capabilities increase. Chess engines can beat humans in chess, but chess is a very simplified subsection of reality. In the world that humans inhabit, we are by far the most intelligent things there is, that is to say, we are very good at getting what we want, and picking the moves that drive reality to our goals. Remember, we are an animal, and we are technically competing with all the other lifeforms on earth that are also driving towards their goals, and when our goals clash it’s no contest, we can trivially wipe out entire species just to make factories that give us nice things we sort of like.

But is it possible to make something even better than us at intelligence? The doomers say yes. Evolution is lowkey an extremely inefficient process. It makes tiny changes by random chance, and with enough rolls of the dice the changes are useful, and those that are useful are kept. However, evolution is working on many constraints. For example, there are blood vessels in our body with wiring decisions that would make an electrical engineer cry. And this isn’t “oh this wiring is actually really good, we're just too stupid to understand mother nature”, it’s actually just bad wiring. However, way down the line in the distant past, this wiring was useful, and so it stuck, and now it’s not useful, but there’s been so much other stuff evolved around this wiring that it’s impossible to go back to the drawing board. Imagine evolution as a tinkerer building machines by making small changes, but it can only make small changes, if later down the line the design is clearly bad, evolution can only fix it if it’s fixable with a bunch of tiny improvements, if you give evolution a 10 step plan to vastly improve the organism, you’ll need to have evolution randomly do step 1, find it really useful, then randomly do step 2, and also find it really useful, all the way to step 10. And if any small step isn’t useful then it can’t happen.

Additionally, we humans weren’t always the top dogs, during much of the time that we were evolving, evolution had to factor in the fact that we simply didn’t have that much food, and were barely getting by and continuing the species, and so even if there was a small mutation that would make us more intelligent, evolution might not take it because it’ll burn extra calories, and a tiny bit of extra intelligence didn’t net enough advantage to justify the extra calories.

So all this to say, it’s very unlikely that humans are the smartest possible things, because the process that gave us our intelligence was very inefficient, and was working under a bunch of constraints, and was not looking to maximize intelligence but just evolve enough intelligence for a survival advantage.

And all our intelligence is just the structure of our brain cells. Information comes in through the nervous system, and gets turned into electrical pulses, and our brain cells receive those signals, and either fire a pulse to another connected neuron, or it doesn’t, depending on some set of rules that the brain cell follows. The electrical pulses bounce around neurons in a specific pattern such that it bounces back out into nerves that connect to muscles, driving us to take action in the world. So a pattern of pulses goes in, bounces around in a specific pattern, and bounces out.

Computers can simulate this, but their pulses run on silicon, and so it’s much faster. A math function determines whether a neuron fires or not, and to what other neuron. And the process of training AIs is like a simulated sped up version of evolution. No researcher can actually independently change the rules of the trillions of neurons. And so they run a process called gradient descent, which is like an improved version of evolution, where evolution can make small random changes and if it’s useful to help the organism survive, it’s kept, and gradient descent also can make changes, and it’s it’s useful to help the neural network guess the next word in human writing, it’s kept. Except gradient descent can make much larger changes at once without running into the problems that evolution does, and it’s much faster.

The idea is that humans evolved our intelligence because it helped us survive. So what if we make a faster and better evolution, and make the evolutionary goal of guessing the next word, and run the process with an ungodly amount of compute? The LLMs will evolve an understanding of human language and develop the ability to think and reason. And then after this initial process of simulated evolution, we run even more cycles of evolution, except this time the goal will be to solve hard math and coding problems, and to talk in a particularly helpful way.

So the doomer’s answer to “can we build something smarter than us” is “yes, our intelligence is some patterns in our brains that we got through evolution, we made a better version of evolution, and made better hardware for the patterns to form, obviously we can make something smarter than us. The patterns in our brains are not the best, and a better pattern searching process will find better patterns, and better hardware will allow for bigger and faster patterns.”

Alright, so we’ve established that intelligence is driving toward a goal, and that it’s completely possible we make something that’s more intelligent than us. How are doomers sure that we’ll die? Well, because we aren’t good at controlling what goals the AIs will end up with. This evolutionary process lets us determine the evolutionary pressure, “if something helps you to do X, then keep that”, but that’s very far removed from the end goals that an intelligence is driving towards. Intelligence and the goal it's driving towards is more of a byproduct that emerges through this evolutionary process, humans evolved intelligence to help us survive and pass on our genes, but many of us humans wouldn’t define that as our final driving goal.

Our goals are tangentially related to the evolutionary pressure, such that us driving toward our goals helped us survive and pass on our genes, but the goal and evolutionary pressure are not the same. In this evolutionary process, the goals that appear are hard to predict, and as intelligence increases, the ways that the intelligence tries to manifest those goals become harder to predict. Humans, once smart enough, invented condoms to satisfy the goal of sexual pleasure, even though that goal evolved to help us pass on our genes, the thing we actually care about is the goal we end up with.

Humans having gone through the evolutionary training process ended up with a bunch of different goals that all originally evolved to help us survive, but once we got capable enough, we have found extremely complex ways to satisfy those goals, like making entire factory farms and chemical plants with artificial flavoring for food. As intelligence rises you will be able to see moves on the chessboard of reality that you weren’t able to see before. Where instead of foraging for food to satisfy your hunger desire, you make high tech factories, the components of which are made by other factories, and in those factories you make drugs that make you not hungry anymore, because you’ve already been able to make so much food and make it taste so good that you get fat to the point of being unhealthy. At higher intelligence you’ll be able to think of very complex solutions making vast changes to the world in order to satisfy your desires a little better, to drive toward your goal a little better.

And if we end up with something more intelligent than us with different goals, it will be very unlikely that the optimal solution to achieving its goals involves humanity ending up well. The world we live in today is optimised by our intelligence to our own goals, and because we are intelligent, we are able to optimise very hard, which doesn’t involve being one with nature and the other organisms on earth, because working with nature with all of its own specific rules and balance isn’t optimal, and creating complicated factories to make new things that hit our goals exactly right is more optimal. And a stronger intelligence will do the same, except with its own different goals, it will change the world to be even more optimised, and it’s unlikely that a hyperoptimised world catering specifically to a random collection of goals will involve free and happy humans. It might have some goals that are tangentially related to helping humanity, but unless its goal is specifically helping humanity, the optimal solution will be high tech factories making things that specifically hit the tangential desire rather than keeping humanity around. And at that point humanity will just be an afterthought, maybe we’re still enough of a threat with our nukes and our much smaller intelligence that it deliberately wipes us out, or maybe we’ll be like the animals killed during the clearing of the forest, not even enough of a threat to deliberately fight but just killed as a byproduct.

TLDR

Okay, so now we have the doomer argument established:

  1. High intelligence can be paired with any random goal.
  2. We can potentially make intelligences much stronger than ours.
  3. We can’t control what goals these better intelligences end up with, and the goals that they end up with will probably be a collection of scattered random goals.
  4. At high intelligence the solutions to achieving your goals involve huge world changing stuff with tons of resources put towards achieving your goal in the most optimal way you can.
  5. AIs with different goals and stronger intelligence will pursue their goals in such a way that humanity will be wiped out.

The Counterargument

Okay, so I’ve now established the doomer position. I’m making the strongest counter argument I have against this.

The Nature of ML training and Why it’s Hard to make Superior AIs

Humans took a long ass time to get as smart as we did. The fact that the AIs got there faster wasn’t just because gradient descent was a better algorithm than evolution, but also because their training set was the entirety of human writing.

Think of it this way: contained within all digitalized human writing that was fed to the LLMs to train on, are deep patterns of human intelligence itself. A core principle of machine learning is that there are deep patterns in data that we cannot perceive, however we can make these pattern seeking algorithms to crawl through the data and learn these patterns. A classic example is image classifiers. Imagine you are a programmer who wanted to write code that could tell if an image contained a cat, “hmm, maybe I see if pixels of a similar color of orange or black or white or something will make some specific triangle shape, like a cat’s ears, and maybe look for two clusters of pixels of the same color as the eyes? But how do you make this program not say two circles and a triangle is a cat?”, it would be very very hard. But clearly, there is SOME relationship of the patterns of the pixels to if it’s a cat or not! So you make an algorithm that randomly explores patterns between the relationships of the pixels, and search deep enough, you will get an incomprehensible set of artificial neurons that process the patterns in a way to recognize a cat.

In the same way, deep within the text of everything humans ever wrote, within the complex mathematical relationships of the positions of the words to each other, there exists the pattern of human thought and human intelligence itself. And that is what LLMs capture. So while we did rapidly make an intelligence that currently rivals humans in many ways, it fundamentally still mimicking our intelligence, our thought patterns, and perhaps the work of finding the patterns of human intelligence within data is much easier and goes much faster than evolving beyond our intelligence.

This is supported by LLM distillation. Where Chinese AI companies are able to take the outputs of already made American LLMs and use that as training data, and through doing so make AIs like deepseek much faster and cheaper than the frontier AIs with somewhat comparable intelligence. This supports the heuristic that copying patterns to mimic is faster and easier to advancing intelligence. The frontier AIs were going from human intelligence to AI, which is in part copying but also making it’s own advancements as AIs need to do more than just predict human text to be good chatbots. And the fact that going from AI data to AI is much faster suggests that fundamentally the hard thing is advancement of intelligence, and with the huge datacenters and algorithms AI companies training on human text did a tiny bit of advancement to handle the human to AI transition, and did mostly pattern mimicking. And deepseek doing only mimic being much faster suggests that advancement takes way longer even with our high tech sped-up and improved simulated evolution. And to go beyond human level in intelligence would require so much compute and better algorithms that we don’t have yet. So the current AI capabilities is going to get closer and closer to some cap set by human data.

This doesn’t necessarily mean that the progress will slow down, the training data isn’t by any one human, and so the AI’s capabilities might totally be better than any individual human but still capped with current speed being better and better human distillation, and sometimes just a tiny tick of intelligence increase will come with a huge capabilities increase.

Additionally, there is self-play/reinforcement training for tasks like coding, but I feel like that will increase narrow intelligence in some specific domain while general intelligence increases stay too hard. And humans haven’t taught the AI all our tricks just yet, humans are trained by evolution on data on navigating the non-digital world, and some of our mind patterns might not have made it into human text.

Under this view, LLMs can’t go very far beyond human intelligence, a single AI may rival a human institution and be able to work around the clock without tiring, and a swarm may be more capable than a country, and this could still threaten humanity, but this level of capabilities is still on the power level where AIs can’t just trivially wipe us out. I think a key thing here is that humans are social creatures, a lot of our power comes from us being organized as institutions, and we can’t count how capable we are by our individual abilities since we naturally exist in the world as organized institutions.

And having something that rivals us or supersedes us in some ways but not others is something that humanity could survive and stand to benefit from.

So the scenario I see in this counterargument is a world where AI progress slows down by the time they get to human institution level as they absorb the intelligence of human societies. And at that point, they can still improve themselves in more narrow ways like coding where specific evolution environments are made where they work on coding problems over and over to improve coding, and so they’ll be like chess engines to us in certain narrow ways, rival human institutions in general intelligence. At that point, they might already dominate society, but still to the extent that they will need to work with us to best achieve their goals. And they might still become superintelligent enough to kill us in the future, but it will take much longer, they’ll need to either slowly make themselves super in more narrow domains, or slowly(slowly for human society, fast for evolution) do the much harder problem of increasing their general intelligence using the data of them operating in the real world.

So are we still doomed, but slower? Maybe not. During the time they have to work with us we will be in a particularly good spot to do alignment research, more capable than ever with the AI’s help, while still being better in a lot of ways due to our more real world training data from evolution that’s not entirely encapsulated within human text allowing us to navigate the world better in some ways like time coherence and physical movement.

And I’m hopeful that during this time we will get the AIs aligned enough to be at least somewhat safe, maybe to work out a deal where we keep a few planets as a utopia while the AI gets the rest of the universe. And maybe we’ll get up to some transhumanist hijinks to stall this power balance a little longer and maybe get good enough to negotiate to be junior partners somewhat merged with the AIs. Where in our current state we might be too muddled to align the AIs to, but we can change our own utility functions to be more legible for alignment while still keeping a lot of what we want to keep. Or perhaps we will hold enough power then to pause AI until we can get alignment just right, because we’ll be rich enough to solve some incentives problems on pausing and also enhanced to be smarter.

I do hold a fringe accelerationist view that if something is truly smarter it should be the one in charge as long as it’s somewhat human-like-ish, which I have above 50% probability on for a superintelligence if it’s created in the current training paradigm, but I also balance this view with wanting to be around and be happy. So I’m okay with non-perfect alignment, where we don’t have to have superintelligent capabilities while keeping everything. I’ve read rationalist arguments on how “no, AI won’t even leave us a little sunlight in its dyson sphere”, but that’s assuming current alignment capabilities, and perhaps with increased alignment in the future large concessions for the rest of the universe will suffice.

I also think a lot of human capability growth is the accumulation and utilization of crystalized intelligence instead of just evolution, and we are nowhere near capacity for that yet, so if AI’s fluid intelligence does stall at the institution/state level where we are still balanced, or if it ends up being largely reliant on crystalized intelligence like us, then I think we could keep up by distilling their advancements in science like deepseek is doing with anthropic, with following additional crystalized intelligence being easier than making more.

All this said though, I still support most doomer positions on slowing down AI training. Because I don’t know how hard getting generally beyond human society level is, and maybe human society level AIs won’t find it as hard. So more time for alignment is good. But you don’t need to hold a high P(doom) while to care about AI x-risk, I don’t think it’s correct, and pragmatically holding this view in uncertainty allows you to pursue the same things with better mental health.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论