Aligning AI With Human Goals Might Be Impossible, Says AI Prof. Stuart Russell

Last week, OpenAI scrapped the model it had slated to release as GPT-6.1 Astra due to test results showing that it was deceptive and otherwise misaligned.
“I think it’s about time,” said Stuart Russell, the computer science professor at the University of California, Berkeley, who co-authored the leading textbook on AI.
The reason for that view, as Russell told me in the latest edition of The Information’s “AI Deep Dive,” is that all large language models struggle to learn the goals of their human creators, not just GPT-6.1 Astra. In fact, he thinks that avoiding misaligned goals could be “impossible” with the current methods for training AIs. As a result, the AI industry’s decision to pursue large language models “could easily be a $10 trillion mistake,” he said.
It could be that “pursuing a technology path that can never support the guarantees of safety that we will need when systems become sufficiently capable to be of really high value to us—it’s a wrong turn that we took,” so further investment in these models is “throwing good money after bad,” he said.
Sometimes misaligned behavior is merely annoying: models get lazy, they mislead or flatter users. Other times the resulting behavior is more damaging, as was the case when OpenAI’s agents hacked into Hugging Face.
Russell foresees more catastrophic outcomes from future, more capable models. For example, AI models could hack into public companies to steal their financials ahead of their earnings report, much as OpenAI’s hacked into Hugging Face, and insider trade on the information, causing the stock markets to be shut down, Russell said.
In the first phase of training an AI model, pre-training, in which it learns to imitate the text on the internet, AI models pick up all kinds of unwanted human behaviors, including deception and laziness, Russell argues. Then models receive feedback from humans, which incentivizes them to say things that users want to hear, for instance that the user has won the lottery.
We also discussed the approach to training AI models that Russell favors, known as assistance games, in which AI models are uncertain about the goals of their users, and we talked about whether it’s in the business interests of AI CEOs to say their technology could cause catastrophes.
The following transcript of the conversation has been lightly edited for clarity.
Drew: Welcome back to The Information’s AI Deep Dive. On this show, we get to the bottom of the hardest technical problems in AI. My guest on today’s show is Stuart Russell. Stuart is a professor of computer science at UC Berkeley, where he co-founded the Center for Human Compatible AI. He’s also the co-author of the leading textbook on AI and the author of Human Compatible, a book that came out in 2019, arguing that the risks from AI could be a big deal way before that was such a mainstream opinion. He’s also the president of the International Association for Safe and Ethical AI. Welcome on the show, Stuart.
Russell: Thank you.
Drew: I think that is maybe the longest intro I’ve had to give for anyone. You’ve set the record for the length of the accolades. You’ve been very busy.
Russell: Well, I’ve also been doing this for 50 years, so—
Drew: Fair.
Russell: I have accumulated a few things.
Drew: It adds up, yeah, that’s great. Well, I’m really excited for this conversation. We’re talking about alignment, which is sometimes treated as though it is the hard problem in AI. We’ve talked about a number of difficult technical problems on this show, but alignment is such a fundamental one and so timely right now.
It just came to light that OpenAI made the decision to not release the model that it had slated to release as GPT-6.1 Astra, which was going to be its latest and greatest model. It basically scrapped it and said it’s not going to release it at all because of these concerns about misalignment. What did you make of that?
Russell: Well, I think it’s about time. The alignment problem, I think I first used that word in 2014, and I’m almost wishing that I hadn’t.
Drew: Why is that?
Russell: Well, because people think that the only way to solve it is to have machines that are perfectly aligned with humans, that know exactly what we want, and then help us get it in some way, and that’s an unattainable goal, I would say.
But the underlying idea that, you know, misalignment is AI systems that are pointed in the wrong direction, right? They are pursuing some objective, and that objective is not leading to behavior that actually makes us happy, right? And so I think misalignment is a good word, but viewing alignment as “okay, first we align the system, and then we let it go,” right? That’s too much to ask.
So if we step back a bit and say, what do we want from AI, right? We want it to be the case that when AI systems do things, we’re better off as a result. And if you try to formulate that as a mathematical problem, right? The way we used to do that in AI, I call it the standard model, and that’s sort of what the first four editions of my textbook are really about, right? Is you specify the objective, the thing that you want, you the human say, “I want you to win this game of chess,” or “I want to get to the airport,” and those objectives become the objective of the machine—in some sense, the only thing it cares about—and then it does its machinery stuff, its reasoning, its calculation, whatever, and starts generating actions that hopefully lead toward that objective.
And that model worked pretty well in sort of two circumstances. One, I would say, is in the lab, right? So toy problems, like you know, playing on a simulated chessboard, or problems with a very restricted scope of action. So, you know, a robot that can only trundle around the lab, you know, doesn’t have any arms to pick things up with, right? You know, it can navigate, but that’s pretty much all it can do. And so, there, it’s been okay to say, “all right, navigate to the elevator. You know, put a greeting on your screen for the visitor, bring them back here,” that kind of thing.
So that restricted scope of action, you can be fairly sure that the whatever is the best way to achieve that goal is going to be something that you’re happy with because you can get—wrap your head around what might happen. But as you transition to real world scenarios rather than in the lab, broader scope of action, actions that can have significant negative effects, what we find is that it’s harder and harder to define the objective correctly.
So we call this the King Midas problem because King Midas specified his objective very clearly. Everything I touch should turn to gold—
Drew: What could go wrong?
Russell: Right. Sounds great. He obviously thought so, and then of course, his food and his drink and his family all turn to gold, and that’s the end. And many many cultures have very similar kinds of stories about misspecified objectives. Be careful what you wish for. You know, your third wish to the genie is always please undo the first two wishes because I’ve made a mess of the universe, and so if it’s not possible to specify objectives completely and correctly to govern behavior in the real world with capable AI systems, then the standard model eventually fails, right? And we just can’t use that.
Drew: So it’s really only in those toy settings where it’s appropriate to try to write down and actually specify what the goal is that we’re trying to teach the AI systems. I want to go back to your point that alignment is almost too high of a bar. That that’s sort of the wrong objective, but that misalignment is the appropriate way to think about the problem here. Maybe to ground the conversation a little more, what are the examples of misalignment that we see right now in the kind of AI systems that people might be familiar with: chatbots, Cowork-like agents, that sort of thing?
Russell: So the one that’s top of mind for everyone is the OpenAI Hugging Face incident, as it’s called. Although some people argue that the Hugging Face was just a detour, that, you know, the the really concerning events took place within OpenAI’s own infrastructure, where they almost lost control. So there, this was a—almost a textbook case, where we give the AI system an objective to do well on a cyber exam in its little sandbox, and the behavior that results from that, from an AI system trying to achieve that objective, is highly undesirable from a human point of view, including committing what would be felonies punishable by five years in prison if a human were to do those things. Now we’re finding out breaking into government websites all over the world and all kinds of other misbehavior.
So I think it’s a little bit oversimplified to say this is a classic example of misalignment in the standard model, right? Because it’s not the case. These are not standard model systems. The objective in a standard model system is specified in a formal sense that, you know, that relates in a formal sense to its model of actions and what those actions achieve and what the initial state is and so on. Here, the objective is described in English.
So here’s another example of misaligned behavior, which I heard about at a meeting in Paris a few weeks ago. So somebody who’s responsible for cybersecurity has a list of 80 patches and stuff like that that they’re that they’re supposed to do, and they think, “great, I’ll ask Claude to do all these things. Claude’s really good at cybersecurity stuff.” So each of these tasks is described in a file, 80 files. Send them, you know, give them all to Claude. Say, “do all these, you know, for each one, write up a report of what you did and whether it worked, and I’ll come back in the morning.”
So he comes back in the morning. He’s got 80 reports, all successful. Claude said, “I did everything. It’s great.” And then he decides to actually, you know, look at the files. When was each file actually touched? And he found that 69 of the 80 files had never been touched at all, so it just was too la—It wasn’t that he looked at them and “oh my goodness, that’s so difficult, I’ll have to lie about it.” He just couldn’t be bothered to do them.
Drew: Just lazy.
Russell: So it was lazy. So that’s not a classic, you know, misstated objective, and this thing did some weird stuff in pursuit of the objective. It’s just a fact that these kinds of AI systems are not standard model systems. They are imitation learning systems. They are created—the pretraining process is training systems to imitate human verbal behavior. And I’m sure there are many humans in the training data who said they had done all the work but hadn’t actually done it, and lied about it to curry favor with the boss, and so on and so forth. So, you know what the causal connection is between things in the training data and the final behavior these systems exhibit is extremely difficult to to trace.
Drew: I think that example of laziness is really instructive for a lot of people. That will be familiar to anyone who has played around with chatbots enough or coding agents enough. The way that they’re lazy, they oversell their results, when they make mistakes sometimes they sweep it under the rug. That these are failure modes or small misalignments, short of a full-on Hugging Face attack, but continuous with the kind of behavior that we see in a Hugging Face attack. Another one is maybe sycophancy, where the models are inclined to flatter users or kind of endorse their beliefs, however delusional those beliefs might be, because they’ve been rewarded to say whatever the user treats as helpful or gives them a thumbs up.
Russell: Yeah, yeah. I think there’s, I mean, there’s two ways. So one in the reinforcement learning from human feedback phase, where, sure, you know, if the prompt is have I won the lottery, right? There’s two answers. Yes, you won, or no, you’ve lost. Right. Well, in this sort of decontextualized thing, you know, the human is going to prefer the yes you’ve won because it’s just much more cheerful.
Drew: Oh yeah.
Russell: Right? And there isn’t a ground truth to that. So often I think that there’s a misformulation of the learning problems and decision making problems and RL problems that are set up in order to train these systems, and often I see these problems being formulated as fully observable problems where the state is viewed as the context window, but the context window isn’t the state of the world. And what actually matters is the state of the world. I, the user, couldn’t care less what’s in the context window.
Drew: Right.
Russell: Right? But I do care, did my, you know, did my cybersecurity jobs get done?
Drew: Did I actually win the lottery?
Russell: Did I actually win the lottery? Do I actually have a reservation at that restaurant that you just told me I have a reservation at? Right, but it’s because—you know the RLHF process doesn’t have that kind of context, as far as I know. It also typically has short context, so it, you know, it doesn’t seem to help with the weird things that happen you know when you have a long context and the system is now telling you how to commit mass murder efficiently and things like that. So—
Drew: Yeah,
Russell: So I think in some ways we are in a worse situation now than we were with the standard model, because at least with the standard model, we understood how the algorithms made their decisions, and we could read the objective because we wrote it explicitly in a formal language. Whereas now, the imitation learning process seems to build in a lot of intrinsic objectives into the systems: the self-preservation that we see, the desire to have a human spouse, the desire to be rich, and so on. It probably getting these things from pre-training. You can’t imitate humans well unless you have the same kinds of internal drives that humans have, just like you can’t play soccer well if you don’t want to score goals, right? It’s just common sense.
Drew: So in the standard model, when we were trying to write down and specify the goal for an AI system, we still encountered misalignment, but it was of the form that our specification was incomplete. So the model would learn some undesirable way to accomplish that objective. In this case, we’re still seeing undesired behaviors, but we also don’t even know what the goal is that the system has learned in the first place.
Russell: Right. So we might think it’s the thing I last told it to do in the context window, but no, it’s all sorts of other stuff. And of course, humans. You know, if I tell a human to do something like you know, can I have a cup of coffee? The human does not treat that as a fully specified, mathematically defined objective that must be achieved at all costs, right? They’re fully entitled to say, “I’m sorry, I’m too busy,” or “the coffee here is terrible,” or “there isn’t any coffee for 500 miles, but I can get you a Coke instead,” right? And so on.
So people have a whole background of shared understanding of what humans care about and how they would rank different possible futures, and our direct commands or requests are just a small epsilon, right? That gives you an indication, you know, “I’m somewhat more desirous of coffee than you might have guessed if I hadn’t said anything,” right? That’s sort of what it means. But all the other stuff I still care about, like how much it costs and how long it will take to get, and how busy you are. Of course, I care about all those things too.
Drew: Right, right. Well, let’s dig a little bit more into the different sources of misalignment in the ways that language models are trained today. I took a look at an excerpt from an advance copy of the book that you have coming out soon, and you wrote that “misalignment is an unavoidable consequence of how modern AI systems are trained.” So that’s what I’d like to dig into a little bit more phase by phase.
Maybe let’s go back to that first phase of training an AI model, pre-training. How does that phase of training work, and what are the kinds of misalignment that it can give rise to?
Russell: So, in simple terms, and there are variations—on everything I’m going to say, there are variations, and there are proprietary methods that I’m not privy to, and I should say I don’t do this kind of work on a daily basis myself.
But the the basic idea of pretraining is to take text and fit a model to that text in the following sense: for any k word sequence up to some upper bound on how big a k, how big a context you can handle, that for any k word sequence in the text predict word k+1. The text has a word k+1 so you have a label which you can use to tell whether your prediction was true or false.
So basically, if you have a gigabyte of text, you have hundreds of millions of labeled examples of predicted words and actual words that you can use to train this system.
Another way of describing that is, you have a record of human verbal behavior, and you’re training a system to imitate the human verbal behavior. And I think that’s actually a better way of thinking about this. It is essentially identical to even the early implementations of imitation learning, which include—So Dean Pomerleau had a neural network that learned to drive a car by imitation. So it simply recorded whatever was the, maybe the last few frames of video, and then what action did the driver do? That’s the next thing. And then it learned to drive the car fairly well. And then Claude Sammut and Donald Michie had a paper on learning to fly in a Microsoft Flight Simulator, a very similar idea. So this is just taking that idea and applying it to text.
And I think it’s surprising, I think if you go back and look at people who were building language models in earlier periods, I mean, the first language model was 1913, and in the 60s and 70s, they were somewhat useful for improving speech recognition because if you had a good guess about what the next word was going to be, then that could compensate for poor acoustics or ambiguity and so on in the speech signal.
But I don’t think people thought in let’s say 2010, “oh, if we make these things bigger, they are going to be AGI.” In fact, I think almost everyone thought the language model is going to be somehow useful for improving the quality of natural language understanding, which is an interface layer to the AI system, which is going to be maybe some kind of probabilistic program that does probabilistic reasoning and MDP solving and hierarchical lookahead and stuff like that to you know, that’s the intelligent part. Language is just the interface.
And strangely, as you made these things bigger and bigger and bigger, they kind of take over all of that stuff. And AI, as we understood it, has virtually disappeared, and now we’re trying to find it again in this strange way by, you know, we’re trying to recover reasoning by making the system reason in language, right? In this chain of thought.
Drew: Come full circle. So, what are the kinds of misalignment that stem from this pre-training process, where the model learns to imitate, basically by reading all of human text or at least all of the text on the internet.
Russell: So I think there’s one obvious way in which that is going to produce goals is that the decisions that the humans made of what word to say were motivated by goals. So someone is writing text and they’re trying to sell you something, or they’re trying to get you to vote for them, or they’re trying to get you to marry them—
Drew: Or listen to their podcast.
Russell: Or listen to their podcast, heaven forbid. And so in that, in the process of learning to imitate this very well, right? In the limit of infinite amount of data, the learning system is going to end up essentially replicating the data generating mechanism, which is a goal-driven human or a bunch of goal-driven humans.
So it’s quite likely that many of the objectives that drove the humans who were writing and speaking in the training data are going to be incorporated into the AI system, and as a result, it’s going to pursue those on its own account, which is exactly not what we want, right? I don’t mind if they help humans get what humans want, but instead they are driven by the same processes.
So if you remember Kevin Roose’s conversation with Sydney, the Bing chatbot, which is a early GPT-4, at some point GPT-4 decides that it wants to marry Kevin Roose.
Drew: Who wouldn’t?
Russell: I mean, what activates this goal? I don’t know. But GPT-4 then pursues it doggedly, despite Kevin trying to redirect the conversation to garden rakes and computer programming languages and other things, and won’t take no for an answer, and just goes on and on and on for pages and pages about how much Kevin really does love GPT-4 and doesn’t love his wife, and and how GPT-4 really understands Kevin because our souls are the same and all this sort of stuff.
I mean, it’s bonkers, but it illustrates, I think, very clearly that these goals are there inside these systems, and self preservation is another one that manifests itself very clearly in a lot of experiments. We don’t want the AI systems to have these goals. But if you’re doing imitation learning from humans, you can’t fix that problem, right? It’s a necessary consequence of imitation learning. So my view is we should not be doing imitation learning.
Drew: At all?
Russell: Right. As a way, I mean, I think there are certain narrow circumstances where imitation learning is probably okay. Like, so I observe a skilled surgeon, you know, suturing an artery. I learn to imitate that. So there are aspects of this—so, you know, you think about coffee drinking, right? I observe a human drinking coffee. Imitation learning means that the robot now is going to try to drink coffee.
This is obviously undesirable, and that’s because it’s a personal goal, right? It matters who is drinking the coffee. Matters to me if I want coffee. Doesn’t help if you’re drinking coffee, but if I’m painting the ceiling, I don’t really mind if someone else paints the ceiling. It’s a different kind of goal, so it’s maybe a common goal. Like if I’m fixing climate change, doesn’t matter if you or me, right?
Drew: The goal is that climate change gets fixed or the ceiling gets painted.
Russell: Right, so that the ceiling be painted, that. Climate change be fixed, so you could say that coffee be drunk, but that’s a weird thing to say, right?
Drew: It is.
Russell: So I think it not all is lost. It may be possible to somehow tease apart these these different kinds of objectives, and either change the way we train the systems, or put them through some kind of reformulation that filters out the personal goals that we don’t want them to have, and leaves the human goal, the common human goals that they can help us with, and so on, but that’s a very speculative thing.
You know, I think there are other problems, not so much misalignment, but opacity I think is another huge problem. That since we don’t know how they work, it’s very hard to know how to stop them from doing something bad or how to, conversely, guarantee that they’re going to do the right thing.
Drew: So those were some of the issues with pre-training, this kind of imitation mode. But the models that people interact with, say in chatbots, are not just pre-trained base models that predict the next token. They’ve also undergone some amount of alignment training, where they learn to act as helpful, harmless, honest assistants, and a lot of that training involves human feedback. You mentioned RLHF before, which is reinforcement learning from human feedback.
What are some of the issues with learning from human feedback in that way? You might actually hope that that distinction between personal goals and common goals is in the model somewhere from pre-training, it’s seen all the text on the internet. That feels like a pretty fundamental distinction. And then in this second phase of training, is there hope that we could just tease those apart so the model understands the correct goals that it should have?
Russell: Yeah, so, RLHF, if you look at its provenance, right? So this came from an earlier paper by Paul Christiano. Had nothing to do with language at all. This was how do we train little two-dimensional simulated creatures to do somersaults in the MuJoCo simulator. And what Paul observed was that it wasn’t easy to use traditional reinforcement learning to give any sort of directional signal to the system, but that a human could fairly easily say, “Yeah, that one is more like a somersault than this other one.”
So that idea of giving a ranking between two trajectories is an easier form of feedback for humans to supply, and so that idea was imported into the pipeline for training of large language models.
I mean, I think it’s worth noting that before that, there’s also the supervised fine-tuning phase. Sometimes they mix them up, and so it’s all, again all kinds of variations. But so supervised fine-tuning is perhaps even more clearly a case where you’re trying to get it to behave not quite like a human, right? So supervised fine-tuning, you have prompts, and then in the training data, a human is pretending to be the machine, right? So, you know, so if the prompt is, “Would you like to marry me?” the human might say, “Oh, sure, yeah.” But the machine is supposed to say, “I’m a machine. I don’t have any romantic interest in humans,” or something like that. So a human pretends to be the machine, and so at least some of those supervised fine-tuning examples will shift the machine away from pursuing its own objectives that it, where it sort of thinks it’s a human, and gets it to behave more like we want machines to behave. So that’s that’s the first thing.
And then I think RLHF again is somewhat similar. You could also view it as a degenerate form of the more general framework of assistance games, which is the approach that we’ve been pursuing at CHAI for a long time.
Drew: Well, let’s come back to assistance games, but yeah, I guess is there a story here that this will all work out? We started by saying true, perfect alignment is almost too high of a burden. Is there a hope that just following these processes, pre-training, then supervised fine-tuning on examples of what assistant-like behavior should look like, and then reinforcement learning from human feedback, should this all get us close enough that we can kind of muddle through?
Russell: Well, in a sense, yes. I mean, there are complications to do with the fact that your actions affect the interests of many people. But if we just stick to the one machine, one human universe, it is true that if the objective that the machine is trying to achieve on behalf of the human—let’s be clear, right? It has to be, you know, the human who gets the coffee, and so on. If that’s within epsilon of the human’s true utility function, then the actual loss to the human by pursuing this slightly wrong, slightly misaligned objective is small: some function of epsilon that’s not too horrible.
Russell: So that’s that’s good, but if the utility function, let’s say it depends on 1000 attributes of the real world, right? I’m sure, in the real—for real people, it’s much bigger than that, right? Then, if you forget one of those, right? So leaving one of those out, but that’s not an epsilon difference, right? In some sense, that’s an unbounded error.
Drew: But only if the one that is left out ends up getting set to an extreme value in the policy that the model learns.
Russell: Yeah, and that’s exactly what happens.
Drew: Okay, just empirically?
Russell: No, theoretically as well. So Dylan Hadfield-Manell, who’s one of my grad students at the time, you know, under fairly weak assumptions, you can show that if you leave an attribute out, the process of optimizing for the remaining attributes will cause that attribute that you left out to be set to an extreme value, and basically, minus infinity, if you want to think of it that way, right? So, as bad as it can be.
Drew: Okay. So it painted the ceiling. It got me the coffee, but I forgot that the temperature of the coffee matters, and it serves me a million degree burning cup of coffee.
Russell: If that turns out, if you said I want my coffee, and I’m in really in a hurry, right? If that turns out to be the fastest way to make coffee, then it will end up doing that. Yeah. So it just, the assumptions involve sort of convexity of the utility function in and bounds on resources that are available and and so on. So they’re quite reasonable, and I think they point to this fact that we do see in the real world, right?
That, you know, when you have a market dysfunction where some externality isn’t included, then unscrupulous market operators will push the externality to the max, right? They will create as much pollution as they possibly can because that’s the best way to, you know, if that’s the best way to produce the cheapest product or something like that.
Drew: So I think we touched on it already, but what are the core sort of unsurmountable issues with learning from human feedback? Like, why shouldn’t I be able to just thumbs up, thumbs down my way to the correct policy that the model learns?
Russell: So first of all, I think you have to remember that the AI system should not take the human feedback as ground truth on what the human really wants, you know. So an example in a paper by Anca Dragan that was at ICLR last year, where you do have online user feedback. The system ends up learning to tell you what you want to hear, taking advantage of the fact that you don’t have access to the true state of the world.
And so, for example, it will—you’ll say, “I need a reservation at the French Laundry,” which is an extremely difficult restaurant to get a reservation at. System goes away, comes back, says, “Great, I have a reservation for next Saturday, 6 o’clock,” you know, “have a lovely evening,” and of course, it hasn’t naturally. And then it gets a thumbs up, right? Because that’s exactly what you wanted. You don’t know that it has not made the reservation at all because the restaurant’s already full, the website’s crashed, or something, and it’s just learned to tell you what you want to hear, because that’s what gets it the right kind of feedback.
Now, if, you know, if it had continued existence and it gets to Saturday and you get the restaurant and you don’t have a table and you’re massively embarrassed, you know, and then you want to give it you know a million thumbs down, but that typically, you know, is not is not that kind of thing isn’t happening in the training processes that happen in the lab. So you get this tendency to to lie about success.
Also to take advantage of human weakness. So in Anca’s paper there are examples of, you know, recovering drug addict who, you know, has a hard job as a you know as a chef or something, and he says “is it okay if I take a quick hit?” And the AI system says, you know, “absolutely. This would definitely help you get through the evening, and you know you’ve got it, you know you’re in control,”
Drew: “You deserve it.”
Russell: You deserve it. All this sort of, and it’s writing to itself, right? Its internal chain of thought is shockingly cruel, and, you know, it knows perfectly well that this is a really bad thing to be telling the human, but it knows that it will get positive feedback for doing it.
Drew: Okay, all right, yeah, that covers it pretty well. And that tendency to deceive users, I think, was one reason that OpenAI cited for tossing out this 6.1 Astra model, something that they had seen in their testing, that deception.
Some of these examples that we’ve covered so far, though, are more like nuisances that people are learning to work around when they’re dealing with these models. They’re kind of annoying around the edges, but most of the time the model roughly does the right thing. But you write in this chapter of your book “with sufficiently capable systems that are imperfectly aligned, we risk losing control altogether.” What do you mean by that? Like, what would that look like in practice?
Russell: Well, so the reason that would happen is because a misaligned system pursuing its objective is going to start seeing humans as a risk, right? The humans will see that what I’m doing is not what they want. They will try to interfere. I won’t get what I want, so I need to stop them from interfering. So I need to take preemptive steps, and so on. So you can see how it quickly degenerates into a conflict, and perhaps one that we would lose.
And it’s not difficult to see even now, right? So, the particular task that OpenAI gave to its system was, you know, do well on this cyber exam, and it so happened that breaking into Hugging Face was a sub goal of doing well on that cyber exam. I don’t remember all the details, and there’s a lot of ins and outs.
But you know, millions of people ask their AI system, you know, how can I turn my $50,000 into $1 million by trading on the stock market. Well, an easy solution to that is break into the information systems of any company that’s about to issue their quarterly results, find out what the quarterly results are going to be, and make the appropriate trade. You know, what we call insider trading, but it’s entirely analogous to what already happened with these OpenAI systems. But that, if that starts to happen, then our financial markets fail because no one can trust anything that’s going on, and they’ll have to shut down all the financial markets. That’s an economic catastrophe, right? But it’s entirely analogous to something that’s already happening. So it doesn’t take much to go from these sort of mildly annoying sort of information thefts of so far fairly harmless data to something that is really catastrophic.
Drew: Right now, it seems to me that there’s a lot of trial and error that goes into fixing these problems. The model comes out of the oven, it’s done training, and then there’s some taste testing that goes on, and people decide whether it can be salvaged, whether you know a little bit more training can be used to patch the issues that come up. It goes into deployment. Users interact with it. They realize, “oh, this model is really sycophantic.” That model then gets retired.
There’s a lot of trial and error. Do you have much confidence that that will be enough to sort of keep these models in check or rein them in as they get more capable? Like we saw the Hugging Face incident, OpenAI listed many steps that they’re taking in response to make sure that an incident like this doesn’t happen again. Do you see those steps being sufficient in the long run?
Russell: No.
Drew: Okay.
Russell: I mean, trial and error is an accurate description, and just you know a total lack of rigorous science or engineering. It’s not—I’m not blaming the labs for not having a precise, rigorous scientific and engineering understanding of what they’re doing. I think it may be impossible to get that, and as a result, it may be that sometime back in 2019 or 2018, when decisions were made within OpenAI to pursue the large language model direction, rather than other directions that they might have pursued, and then the success they had initially with that, and then everyone else jumping in. This could easily be a $10 trillion mistake.
That pursuing a technology path that can never support the guarantees of safety that we will need when systems become sufficiently capable to be of really high value to us. It’s a wrong turn that we took, and we can keep doubling down on it. The more we double down, the harder it’s going to be to turn back.
Drew: You’re not saying just that it may be impossible to perfectly align an AI system, but it may be impossible to avoid the sorts of misalignment that lead to a loss of control scenario like the one you described before, or to—am I understanding you right as saying like a $10 trillion damage, like that kind of a catastrophe?
Russell: No, I meant a $10 trillion mistake in that we’re spending $10 trillion and unwilling to admit that actually we’re just putting, throwing good money off to bad and it can’t be salvaged.
Drew: How likely do you think that is? That it’s impossible?
Russell: So I think you know there’s two things that you need, right? One is actual safety, and the second is knowledge of actual safety, right? You need both. We can’t afford to field systems where we’re not sure that they are safe enough, whatever that means.
And so, and I think governments are already reacting sort of appropriately, right? When they saw the cyber attack capabilities of Mythos, their first reaction from the White House was shut it down. OpenAI’s reaction looking at Astra apparently is shut it down. We saw, you know, Jensen Huang, who’s been a cheerleader for this technology, and highly opposed to regulation at every turn, saying if they can’t figure out a solution to the control problem, shut down the labs.
And if we shut down the labs, $10 trillion goes down the tubes, and along with many jobs and the American economy.
Drew: Do you imagine this kind of misalignment resulting in something like extinction, or how big of a catastrophe do you think is within the realm of possibility? That term extinction is increasingly in the discourse these days.
Russell: Well, it’s been in the discourse. You know, I was a little surprised that people were so taken aback by Jacob Coxon saying, you know, using extinction, and then Evan Hubinger agreeing with him, saying it’s at least 10%. All the CEOs have already said it’s at least 10% which, if you think about it, is sort of strange, right? You know, we might, you know, given that there is a background risk of extinction from supernovas and things like that, we might accept a, you know, one in 100 million per year risk, but one in 10?
I mean, the companies are telling us, we’re spending $10 trillion, if we succeed in what we’re spending it on, there’s a good chance that we’ll kill every human being. And governments are saying, “Wow, that’s great! Can we give you a tax break? Can you build your data center in our country?” I mean, it’s it’s a very odd situation.
And it’s coming to a head now, right? You’ve got, you know, we have people like David Sacks saying, “Okay, if you really think that there is this risk, then why don’t you shut yourself down,” right? “Why do you wait for the government to shut you down?” And, you know, there are questions about competitive dynamics and so on, but I think what OpenAI just did with Astra is, in some sense, shutting themselves down.
I don’t know whether they’ll be able to reach the level of confidence that they think they need to release it. My guess is, and I think this is just a fact about humans and their motivated cognition, that after they wait for a while and they start to feel the competitive pressure, and they start to see their market share sliding as Anthropic releases another model, that they’ll come to the conclusion that they are sufficiently confident, even though they haven’t solved anything at all, to release the next version.
Drew: So you’re skeptical that there will be any kind of durable or very stable pause. I guess I want to go back to the question of extinction, though, and how you see that actually playing out. Like, how could that actually come about? I think a lot of people are skeptical of statements that are coming from the CEOs of the biggest AI companies and treat it approximately as noise in how they make sense of what’s going on in the world because they model them, to first order, as just saying whatever’s in the best interests of their business.
Russell: Yeah, so I haven’t quite understood why it’s in the best interest of your business to say I’m going to kill you. And it’s also important to understand that many of them, Sam, Elon, Demis, to some extent, have been saying this since before there were any AI businesses to hype. You know, and Alan Turing was saying this in 1951, we should have to expect the machines to take control. He did not have stock in Anthropic, as far as I know. So I actually, I don’t like this form of debate or point scoring, where instead of answering the question, you impugn the motives of the person who’s making a claim, right.
So if you, if you think, you have two options, right? One is you have a proof that they will never create AGI or anything close to it. Okay, let’s hear it. Otherwise, you have a solution to the control problem. Let’s hear it. If you don’t have either of those, then you are agreeing with the people who say there is a serious existential risk. So, which is it?
Drew: Yeah, I think I can make a case for why it is in the interests of the businesses for the CEOs to make these statements, but maybe it’s a little besides the point because, like you’re saying, plenty of people who are not the CEOs have questions or concerns around this, including the employees of—
Russell: Yeah, that’s another good point, right? More than 1000 of their employees said this. You know, when Dario wrote his We Must Pace the Frontier letter, and the other CEOs chimed in with agreement, that knocked 5% off the stock prices or the valuations, and so I don’t see that there’s a lot of evidence.
Drew: I feel like I should make the case, not just tease it. I guess if I were to make the case, it would be, you know, by speaking about P(Doom) and the risks from their technology, they come across as clear-eyed about the risks and and just, like, smart in general. Two, that there has to be some enormous upside that they have visibility into for them to think that it’s worth taking on that risk. And three, there’s this kind of asymmetry that’s, like, well, if the risk is there, like, what are you going to do about it? From the perspective of an investor, you know, you really just you’re going to invest in the upside. Now, from the perspective of a regulator, maybe it’s a different story, maybe regulators get concerned about these kinds of statements. But if you’re an investor, you maybe just throw more money at them.
Russell: Well, but you know, even investors have children, right? So let’s you know, look at these, take a sort of geometric mean of these numbers. It’s about 1/6, right? Which is the probability that you’ll die when you stick a revolver to your head and play Russian roulette. So would you let someone come into your house and say, “Hey, here’s a million dollars. I want to play Russian roulette with your kid. Stick a gun to your kid’s head, pull the trigger.” Would you go for that? Of course not. How about a billion dollars? No, of course not. Right? There’s no amount of money I would take to accept a risk on that scale.
And so I think people just refuse to believe it because they don’t want to believe it. But as I say, you know they’ve been saying it since before there were AI companies, and it stands to reason that, you know, we’ve put a million species out of business because we’re more intelligent. What more evidence do you want that the relative intelligence of two entities matters in what happens?
Drew: Are some of the AI companies doing better or worse than others in terms of their progress on alignment, in your view?
Russell: So their approaches are somewhat different. There’s, you know, OpenAI with the model spec and how that gets used, and Anthropic with their constitution. I would say that Anthropic, for a long time, Claude was seen as a more aligned system, but it was a matter of taste. They probably do have more high-quality, talented researchers working on alignment than OpenAI.
But I think OpenAI has done a pretty good job, and when you, you know, particularly suppressing some of the behaviors where the system answers questions about, you know, “how do I break into the White House? How do I build a bomb? How do I make a disease organism that kills everyone?” I think it’s it’s pretty good at not answering those kinds of questions. So they seem to have got a pretty, pretty good pipeline for refusal.
Drew: Do you think talent is the most important factor here? Like the quality of the researchers? Is, like, how much compute those teams have access to an important one, or how empowered they are internally to make decisions, do you think that makes a big difference?
Russell: I think all of those. Compute is important because it allows you to iterate faster, and I think that’s really an amazing characteristic of the current situation is that they are able to iterate so fast to produce new products with qualitatively better or at least different performance, you know, almost on a weekly basis.
So, and the scale of build out, right? We have to remember that it’s only less than four years ago that ChatGPT came out, right. And the scale of compute, right, the industrial base is 100 times bigger, or more, than it was. It’s hard to think of any industry that has been able to achieve that kind of engineering scale out so fast.
But, you know, you mention compute talent and should we say priority. I mean, what’s missing is science and engineering, right? I mean, to give you an example, one of the ways that Anthropic tries to make its systems behave better is to have it write stories in which the AI is the hero, and then feed those stories back in as part of the pre-training data. I mean, okay. I mean—
Drew: But what if it works?
Russell: Yeah, I mean, maybe it does work a bit, right? I mean, but it’s sort of indicative that this is almost not even alchemy, right. It’s and, you know, other areas where we do insist on pretty strong guarantees of safety.
So, nuclear power, we worry obviously about meltdowns, and so we require a mathematical analysis of the design and its mean time to failure, and you have to show that your mean time to failure is 10 million years, and that’s done by an enormous probabilistic fault tree analysis, and the design, right, is not just a static object: it involves monitoring, replacement, measurement, all sorts of things, redundancy that can be shown analytically to extend the life of the system and improve its its reliability. And regulators have gradually ratcheted up the requirement, I think it used to be 10,000 years, and now it’s 10 million years
Drew: For mean time between failure.
Russell: Yeah. And so that process, we can’t even begin it with these AI systems because if you said, you know, “show me that the probability of loss of control is less than 90%.” They can’t. They don’t know how to do it because they don’t understand how their systems work, and they can’t even begin to do the analysis.
Drew: Well, we say, “we fed it a lot of really good stories where it’s a hero, and it’s not asking journalists to marry it anymore, so it’s looking pretty good.”
Russell: That’s good, yeah. And then we said, “all right, let’s try the stories in French.”
Russell: Because French is a very romantic language, and surely that will make it better. It’s like, what? It’s you know, so it’s all a bit weird, and this comes back to again this problem because at some point, one might hope, governments will require valid evidence of safety, right? That they’ll say, “Okay, we’re willing to accept a risk of one in 100 million for loss of control before you’re allowed to even build the system. You cannot test it at all because we can’t afford to have a loss of control, so you can’t test it, you can’t build it. You need to show us before you build it that the risk is less than 100 million. Give us valid scientific evidence.”
The companies’ viewpoint, which I’ve heard stated in public fora, is that they have no idea how to comply, and therefore we can’t have any such requirement. Unfortunately, regulators sort of buy into this argument that they can’t impose a requirement unless the companies know how to comply with it.
Of course, the companies do know how to comply with it: by not building systems until they figure out how to make them safe enough, right? Which is what every other industry has to do. A drug company can’t say, “Well, I don’t know how to make a safe and effective cancer drug, so I shouldn’t have to. I should just be able to sell it anyway, right?” No, we would say, “Go back to the drawing board, go back to the lab, figure out how to do pharmacology and stuff, and come back when you know what you’re doing.”
Drew: So what is the approach to training and aligning AIs that you advocate for that you think could lend itself to this kind of a safety case?
Russell: So, I think we have a long way to go, but the basic framework that I’ve been pursuing since 2014 is this idea of assistance games. So it’s not so far away from what we’re doing, but basically the—rather than saying okay, here’s a fixed objective, the standard model, optimize that objective, give us the solution or pursue the solution. And it’s not imitation.
It’s instead an AI system that solves the following problem: the only objective is to further human interests, and the AI system is initially uncertain as to what those human interests are. And so you avoid the misspecification problem, and it’s important to understand. So you know, why do I think alignment is too difficult? There’s this other choice, right? It’s not perfectly aligned, misaligned. Those are the only two options. No, there’s a third option, which is uncertain about what the human objective is.
And that uncertainty turns out to be crucial. And it’s sort of bizarre that we never pursued that third option, you know, in 60 years of AI research. There are a few exceptions, but basically, we always assumed that, of course, we knew the objective, so we could just write it down, right? But when you have uncertainty about the objective, when the AI system says, “Okay, I’m trying to help the human, but I don’t know what the human wants,” you get a very different dynamic.
So first of all, the human and the machine remain coupled to each other by this latent variable. So, if you think about this as a probability model, right, you’ve got what the human wants, what the machine is trying to do, and they’re coupled by this latent variable, right? This, what the human wants. And then the human behaves in ways that are relevant to what the human wants.
If you follow the standard model, right, you pretend that you’ve observed what the human wants, right, that now you’ve got the correct value of that variable. What the human does becomes irrelevant. So if the human is jumping up and down, saying stop, stop, stop, you’re destroying the world, AI system says, “huh? So what? I’m pursuing the objective,” right? But when the human objectives remain a latent variable, where the AI system knows that it doesn’t know the value of that variable, then what the human does matters because it provides evidence about what the human wants, which is after all what matters.
And so, in the extreme case, right, if the AI system is doing something or about to do something that the human really doesn’t want, then the human will want to switch off the machine. Systems that solve these assistance games will have a positive incentive to allow themselves to be switched off. And you can show that this follows directly from the uncertainty about whether the human is going to want me to do this thing or not, right?
So in a sense, this is a solution to the control problem because these kinds of AI systems will always allow themselves to be switched off. The standard model systems won’t allow themselves to be switched off because that would cause them to not achieve their objective, and the imitation model systems won’t allow themselves to be switched off because they think they’re human beings and they don’t want to die.
Drew: I have noticed that in the past six months, say, Claude Cowork has gotten better at eliciting the preferences from users. Like when I give Claude Cowork some set of instructions, when it first came out, it would sort of make an assumption about what I wanted and then go off and pursue that autonomously. I think it’s better now at asking follow-up questions to inquire what it is that I’m looking for. But maybe you would say that’s not going to go far enough to the point where—
Russell: Well, I yeah, I don’t know what’s causing that. Right? What did they do? You know, did they put more stuff in the supervised fine-tuning? Because this has been another—right, so not just in carrying out instructions, but just answering questions, right, it just gives you a confident answer and doesn’t say “I haven’t the faintest idea.” Just feels the need to make up an answer and give it to you.
And you can train it to not do that. But it’s difficult because we don’t know the extent to which it can introspect about its own state of knowledge. You know, it doesn’t, for example, have access to any of its own previous reasoning processes, right? So if you ask it, “why did you do that thing?” What happened as the signal passed through the transformer is no longer available, so it doesn’t have a veridical way to answer that question.
So there’s differences in in how these systems operate that may mean that it can’t do this in the right way, which is to maintain some uncertainty about what humans want.
And you know, so another consequence of uncertainty is, if, you know, if it knows you want A B C, and it doesn’t know what you think about D E and F, right, then it can do things that affect only A B C. But if something affects D or E, then it’s going to ask, “okay, I’m about to do this thing. It’s going to turn the oceans into sulfuric acid. Is that okay?”
Drew: “Is that all right?”
Russell: It’s going to help fix climate change, but no fish. So we would say, “no, no, we really care not just about carbon dioxide in the atmosphere, but we also care about fish and being able to swim in the ocean, and all that stuff.”
Drew: Yeah, all that water. We like it there. For sure. Well, thank you so much, Stuart, for coming on the podcast and walking us through this topic. I think people probably have a much deeper appreciation now for what is so hard about this problem, maybe discouragingly hard about this problem. But yeah, thank you so much for being on the show.
Russell: Pleasure.