What just happened? Pragmatism and Pessimization

This post is about the major role alignment researchers played in advancing the frontier of AI capabilities over the last decade, and how the distinction between “alignment research” and “capabilities research" thereby lost most of its meaning. In particular, I’ll chronicle the development of what I’ll call the “pragmatic alignment” paradigm, and how it helped the three leading AGI companies push hard on the path to AGI under the banner of safety. This was not a subtle effect—it’s apparent even to outsiders who investigate the field, like authors Sebastian Mallaby and Karen Hao.

In my previous post, I summarized the alignment community’s plan as “differentially advancing alignment over capabilities”. However, it’s worth being more precise about who was nominally pursuing that plan, because it doesn’t seem to have been very action-guiding for MIRI. For example, in 2015 Nate Soares described MIRI’s “deconfusion” research as being guided by the question “what would we still be unable to solve, even if the challenge were far simpler?”. Meanwhile Eliezer’s author surrogate in this 2018 post repeatedly emphasizes that people shouldn't draw direct links from MIRI’s research to its potential applications. So my sense is that the “differential impact” criterion started off as merely a background consideration, then became much more load-bearing with the rise of EA-style thinking in the field, which involved justifying research directions by appealing fairly directly to their consequences.

This made people less rational both on an individual level and on a group level. On an individual level: it’s easy to generate rationalizations for why a given line of research is impactful on the margin, because there are many possible scenarios for how the future could play out (or how the past could have played out if you hadn’t intervened). So external incentives (or even just a strong emotional drive to have impact) can easily lead you to focus on the possibilities which suit you best. Especially within AGI companies, this gave rise to extremely motivated reasoning about counterfactuals in which alignment-branded interventions didn’t happen, helping people deceive themselves and others about their actual motivations. More generally, “differentially advancing alignment” is hard to demarcate from other consequentialist goals like “preventing overhangs” or “buying more time for alignment work later” which leave even more room for deception.

The group-level problem: the alignment community was very bad at dealing with these adversarial dynamics, which meant that it wasn’t able to prevent the gradual erosion of the boundary between alignment and capabilities research. In particular, extreme fear of publicly criticizing powerful people—and strong charitability/mistake theory norms—prevented the community from creating common knowledge of who was doing motivated reasoning, or just straightforwardly lying. Even people pursuing enormously power-seeking strategies—most notably Sam Altman and Dario Amodei—were given the benefit of the doubt for many years. What I mean by “pragmatic alignment”, then, is the whole complex of people who were using and accepting consequentialist arguments about how to make AGI go well, while being emotionally and strategically committed to almost never calling out misuse of those arguments. 

Pragmatic alignment is just one facet of the community’s unwillingness to directly challenge power structures, most notably exemplified in its lack of criticism of AGI companies until recently. To get a sense for how deep-rooted this resistance was, it’s worth reviewing the comments on this post by Ben Hoffman, and this post by Adam Shimi. (There were very few other discussions of this topic before ChatGPT; the most notable are Scott Alexander’s original objection to OpenAI and Jacob Hilton's partial defense of OpenAI.) I should add that, upon revisiting Adam’s post just now, I found that I’d strong-downvoted it—I think because, when I first read it, I was scared of the alignment community alienating OpenAI. I feel quite viscerally horrified by this reminder of how sycophantic my past self was.

More on that in later posts. This post will focus specifically on a historical analysis of how the concept of “alignment research” was twisted towards boosting capabilities at OpenAI, DeepMind, and Anthropic. As I stated in my previous post, the most important point here is not that I’m confident that accelerating AI capabilities has been bad for the world—that would require a level of large-scale consequentialist reasoning which I can’t do reliably. However, what I am confident about is that people who tend to produce the opposite of their stated goals (a process I call pessimization) can’t be trusted with great power, and communities that fail to hold them accountable also can’t be trusted with great power.

By “hold accountable” I’m not referring to any centralized judgement process—we don’t have institutions reliable enough for that. Instead, I want individuals (like you!) to demand honest public conversations about what happened and what should have happened. People’s willingness to have those conversations (and your personal evaluations of how sincere they are) should then guide your decisions about who to work for, who to fund, and who to affiliate with more generally. At this point, someone in the field merely being open to alternatives to the current failed paradigm is sufficient to make me feel solidarity with them. Unfortunately, such openness is often constrained on an emotional level by the desire to remain part of existing networks of power, money, and ideological security.

I also want to be very clear that I’m trying to hold alignment researchers accountable not because I think they’re less ethical than other elite groups, but rather the opposite. Alignment researchers (especially the ones who have been around since the early days) think about their impact on the world more seriously and earnestly than any other comparably-sized community. (By contrast, we should interpret almost every capabilities researcher as being steered primarily by local incentives and power gradients, in a way that psychologically prevents them from seriously considering unconventional strategies.) This gives me hope that (some subset of) the current field of alignment is able to learn from its mistakes. The first step is acknowledging that there’s no longer any widespread (implicit or explicit) definition under which “alignment research” (let alone “AI safety”) is robustly good for the world, based on the evidence I lay out below. The second is adopting more defensible norms and accountability mechanisms, like the ones I discuss at the end of this post.

The Prosaic Ideal, the Pragmatic Reality

Around the time that OpenAI was founded and OpenPhil became active in the field, alignment started undergoing a partial paradigm shift towards a more pragmatic and empirical approach. One early step was the Concrete Problems in AI Safety paper (which I’ll discuss in more detail in my next post). Dario Amodei was both the lead author on this paper and one of the main people formulating this new approach. However, he was still new to the field, and didn’t write much publicly about his views (though this 2014 discussion with Eliezer is a useful source). My impression is that Carl Shulman had some similar ideas but also didn’t articulate them publicly until significantly later. So I’ll focus on cataloguing the shift with reference to Paul Christiano’s extensive public writings, which were the main intellectual arguments updating the alignment community’s worldview. Note that I'm grateful to Paul for recording his thinking in enough detail that I can try to trace what went wrong a decade later; readers should keep in mind that many others influenced the events I describe in less legible ways that make accountability harder.

Paul had been active on LessWrong since 2010, and had started doing significant alignment research by 2013. In addition to authoring several agent foundations papers, he blogged on a wide range of topics. In the following years he developed a new perspective on alignment. The most concrete milestone was his 2018 post arguing that we’d see a slow takeoff; another was his 2019 post articulating more gradual threat models than Yudkowsky’s. In some ways, these posts built on Hanson’s side of the Hanson-Yudkowsky foom debate, but Paul was more willing to accept the premise that general intelligence would be a really big deal, and merely dispute the trajectory by which we would reach superintelligence. In hindsight, he has been vindicated in his arguments for a much slower takeoff than Eliezer originally predicted.

Another important part of Paul’s new paradigm was the idea of "prosaic AGI": an AGI built in a way “which doesn’t reveal any fundamentally new ideas about the nature of intelligence or turn up any ‘unknown unknowns.’” In a sense, the whole field of deep learning is a prosaic approach to AGI, compared with previous methods. But even after its early successes, the additional belief that deep learning would scale up easily took longer to propagate. Dario Amodei wrote a long, never-released google doc advocating for the “big blob of compute” hypothesis around 2018. The publicly-available posts with the most similar content are probably Sutton’s bitter lesson post and Gwern’s scaling hypothesis post. I also recall Jan Leike giving a presentation to the safety team at DeepMind in 2019 arguing for ~6-year timelines based on similar intuitions. My sense is that almost nobody else at DeepMind except Shane Legg was sympathetic to this view.

The prosaic AGI intuition has been vindicated since then: we’re now much closer to building AGI, and we haven’t learned any fundamentally new things about intelligence in the process. But the reason I only called it a partial paradigm shift is that Paul didn't manage to carve out a defensible research strategy. His original prosaic AI alignment post argued against two separate camps. On one side, he critiqued people who claimed that "it’s impossible to do meaningful work without knowing more about what powerful AI will look like". This reasoning is similar to the arguments Dario and Geoffrey gave for working on scaling up LLMs. On the other side, he critiqued people who claimed that "aligning prosaic AGI is probably infeasible". My understanding is that MIRI used this claim to justify trying to build (agent-foundations-based) AGI themselves (see Wei Dai's comment on my previous post for more details).

So Paul seems to have been trying to steer a path between two opposing “alignment” strategies which both prescribed building AGI yourself—an admirable intention, if so. My diagnosis is that he didn't succeed because he made versions of both the individual-level mistake and the group-level mistake that I described above. The former involved characterizing prosaic AI alignment as being in opposition to "understanding intelligence". My sense is that both MIRI and Paul were implicitly treating “understanding intelligence” as mainly valuable for building aligned AGI from scratch—which wouldn’t count as prosaic AI alignment. However, there’s another possibility: that an AGI which would otherwise be built without an understanding of intelligence is aligned using an understanding of intelligence! I’m currently excited about agent foundations precisely as a strategy for aligning otherwise-prosaic neural-network-based AGIs—but this strategy is implicitly ruled out by Paul’s framework.

This mistake was exacerbated by Paul’s strategic mistake of joining OpenAI to work on the same projects that other people were justifying for very different reasons. Because of this, the success of Paul’s empirical predictions (and his general thoughtfulness about alignment) was then taken as evidence in favor of OpenAI’s research directions and overall strategy. Paul conspicuously failed to correct this impression by critiquing OpenAI publicly—I can’t find any critical statements from when he worked there, only an endorsement of the OpenAI safety team (which was run by Dario). It's very normal not to publicly criticize your boss or your company, but for anyone who’s trying to significantly influence the world—and especially a leader of a key movement—the willingness to do so seems like a very basic foundation for maintaining integrity. In the absence of that, Paul’s “prosaic AI alignment” paradigm devolved into a paradigm in which essentially any consequentialist arguments for building AI systems or allying with AI companies were accepted as valid AI safety strategies, as I’ll catalogue in the next three sections.

OpenAI

The intermediate step between “actually trying to solve the alignment problem” and the fully-pragmatic paradigm was scalable oversight. Around 2018, three maybe-probably-equivalent scalable oversight proposals were floating around: Paul’s iterated amplification, Geoffrey Irving’s debate, and Jan Leike’s recursive reward modeling (Jan started at DeepMind, but moved to OpenAI in 2021). Iterated amplification was by far the most-discussed amongst alignment researchers. Paul’s (notoriously opaque) arguments focused on the idea that if imitating humans is safe, then we can combine many imitation learners to produce more capable (but still safe) agents. However, I broadly agree with Yudkowsky’s critique that this hides the hard part of the problem in the interactions between the subagents.

More importantly, whatever theoretical merits these proposals had were immediately decoupled from the engineering work that Paul, Geoffrey, Jan and Dario actually started doing—specifically, work on reinforcement learning from human feedback. This started with agents learning simple behaviors in toy environments, but soon progressed to a series of papers applying RLHF to LLMs, culminating in InstructGPT. While these were impressive efforts on an engineering level, there’s very little that distinguishes them from what a prescient capabilities-maximizer would have been doing—for example, although Paul's theoretical justifications for iterated amplification referred a lot to the safety properties of imitation learning, all of these papers added RLHF for better performance.

This focus on engineering-style work was facilitated by Dario’s push to scale up from GPT-1 (which was mainly Alec Radford and Ilya Sutskever’s project) to GPT-2 and subsequently GPT-3, justifying this in significant part by arguing that it would help boost alignment research. In Empire of AI, Karen Hao reports Dario telling her in 2019 that “We want a language model that humans can give feedback on and interact with [where] the language model is strong enough that we can really have a meaningful conversation about human values and preferences.” My understanding is that Paul opposed this strategy internally, but Geoffrey supported it. The Infinity Machine quotes Geoffrey as recounting ““We struggled for a while [to get LLMs to obey instructions]. Then we were like, OK, let’s just make the language models stronger.”

Subsequently, Dario led the effort to scale up GPT-3 training to 10,000 V100 GPUs. In addition to arguments that better safety research required more capable models, my understanding is that he was also trying to increase OpenAI’s lead against China; I’ll discuss that kind of reasoning in more detail in a later post. Before leaving OpenAI, Dario also released the scaling laws paper, which did a lot to wake the academic ML community up to the plausibility of AGI. I don’t have a strong sense of what we should infer from this, since I’m predisposed to be positive about scientific communication, but it does seem like more evidence against the idea that Dario was following a coherent and sensible plan.

Meanwhile, John Schulman had been at OpenAI from the beginning. He was sympathetic enough to safety to coauthor the Concrete Problems paper, but primarily worked on reinforcement learning (e.g. pioneering PPO). By 2021 (the year I joined OpenAI) he was working on WebGPT, a way of letting GPT models browse the internet. I remember him articulating reasons to think of WebGPT as an alignment project (something like: if models can look up information online, they’ll be more honest). These justifications were apparently sufficient to get a number of alignment-motivated researchers to work on it (in particular Jacob Hilton—the first author of the blog post—Jeff Wu, and William Saunders). WebGPT was the direct predecessor to ChatGPT, and my understanding is that ChatGPT inherited a lot of WebGPT’s codebase (as well as ideas and techniques from InstructGPT). More specifically, a researcher who was on the team around that time described ChatGPT to me as “WebGPT minus the Web”: the basic Q&A format and RLHF fine-tuning were already there, but ChatGPT lacked WebGPT’s unreliable web browsing component.

Paul has since written up his justifications for working on RLHF, and why he doesn’t think RLHF was very important for ChatGPT. However, these arguments seem very suspect (for reasons explained well by Habryka). For example, Paul says “I think the effect [of ChatGPT] would have been very similar if it had been trained via supervised learning on good dialogs”. But the InstructGPT blog post reports that “our labelers prefer outputs from our 1.3B InstructGPT model over outputs from a 175B GPT‑3 model [trained with supervised fine-tuning, as per Figure 1 from the paper], despite having more than 100x fewer parameters”. Another important datapoint comes from Sydney Bing, which wasn’t trained with RLHF and produced fairly unhinged outputs, suggesting that RLHF was important for making ChatGPT user-friendly.

In hindsight, the launch of ChatGPT was one of the most acceleratory events in the history of AI, funneling many billions of dollars into the field (ChatGPT grew faster than any previous product in history). I don’t have a great recollection of whether or how the ChatGPT team justified this launch in safety terms; I expect their reasoning was that someone else would do it if they didn’t. But as I’ll discuss shortly, the main potential “someone else”s were also researchers nominally motivated by alignment, who were also justifying their work with the idea that someone else would do it anyway. At the very least this was a colossal coordination failure within the community; I also think it undermines the core premises people were using to reason about how to have impact.

One such premise was the idea that, if a system was developed using relatively few resources, it could likely be quickly scaled up to many more resources, which might create a dangerously sharp transition. The possibility of such “overhangs” was discussed at least as far back as the 2008 Eliezer-Hanson debate (with hardware as the limiting resource), but only as a background strategic consideration. At some point, people started using overhangs as justification for making rapid progress now, to use up all the low-hanging fruit so that later progress would be slower (and therefore less dangerous).

Reasoning of the form “we’ll do something we’re worried about so that other people do less of it later” is always extremely slippery, in a way that common-sense morality (and even just common sense) weighs strongly against. This case was no different. Broadly speaking, people would pick whichever inputs to AI progress they wanted to defend speeding up, and just assume (often even without directly stating it) that there were other background constraints which meant that speeding up their preferred inputs wouldn’t make much long-term difference.

In case this seems like an exaggeration, consider these two discussions of overhangs from Paul:

“If LM agents are weak are due to exceptionally low investment and understanding it creates "dry tinder:" as incentives rise that investment will quickly rise and so low-hanging fruit will be picked. While there is some dependence on serial time, I think that increased LM investment now will significantly slow down progress later.”

And from this post:

“Avoiding RLHF at best introduces an important overhang: people will implicitly underestimate the capabilities of AI systems for longer, slowing progress now but leading to faster and more abrupt change later as people realize they’ve been wrong. Similarly, to the extent you successfully slow scaling, you are then in for faster scaling later from a lower initial amount of spending—I think it’s significantly better to have a world where TAI training runs cost $10 billion than a world where they cost $1 billion.”

The most obvious, basic model of progress is that things take time, so doing stuff earlier will allow people to do more stuff later. Indeed, one of Paul’s most significant intellectual contributions was the argument that recursive self-improvement will continuously ramp up over time—which implies that pushing AI forward will have compounding effects. It’s possible in principle that local bottlenecks could override these dynamics. But if we argue for speeding up algorithmic progress and investment and public understanding (and even elicitation) of AI capabilities based on overhang arguments, then there’s almost no room left for limiting factors to kick in later. There’s something like a “bottleneck of the gaps” here—i.e. the “limiting factor” is whatever some safety person hasn’t decided to work on yet, and tends to zero as AI safety people find arguments for accelerating every possible input to AI capabilities. (What about the difficulty of getting US visas for AI researchers? Remco Zwetsloot and other DC safety advocates have worked on it (see section 5.1). What about the fact that Europe isn’t a leading player? I’ve talked to several AI governance people who are considering kickstarting a Europe-wide AI project, on the grounds that they like European values. And so on.)

The strongest fallback for overhang advocates was the difficulty of increasing the hardware supply. However, these arguments are also looking very shaky. AI progress has redirected capital at a civilizational scale: AI investment was 39% of US real GDP growth in the first nine months of 2025, and “capital expenditure of just five technology companies is now larger than global investment in oil and natural gas production”. As a cynic would expect, AI safety people have specifically been homing in on the highest-leverage investments—for example, Situational Awareness just invested $400 million to disrupt a key chip production bottleneck. In some sense the “compute overhang” argument remains unfalsifiable, because we can always construct counterfactuals which are worse than our current situation. But for any practical purpose, the final nail in its coffin is the fact that so many AI safety people are now taking seriously the idea of an imminent “software-only singularity”—see Tom Davidson, Ryan Greenblatt, and Paul himself (in non-public talks and writing). Insofar as they’re right, all work which was (explicitly or implicitly) justified by the idea of reducing the hardware overhang has been directly pulling us towards the singularity. (To be clear, I don’t expect a software-only singularity; my point is that the worldview which accepted “overhang” justifications is no longer coherent.)

What went wrong here? It’s hard to know exactly what led any given person to endorse any given argument. But when we zoom out, it becomes clear that many people in this space really want to pull some lever that feels important, and privilege arguments which justify that. That might come from a sense that they need to have an impact on the world; or fear about failing to fulfil their potential; or the more mundane explanation that big levers tend to be associated with money and prestige and proximity to power. Certainly the latter was a large part of why I joined OpenAI originally; I expect that most people who joined earlier were less prestige-oriented than me, but still made that decision using reasoning that was warped by similar emotional drives (and later further warped by the social dynamics of actually working there).

To describe that warping, I find a version of Ajeya’s saints, sycophants, schemers trichotomy useful (though I think of it as a spectrum between fully scheming and fully sincere). Ajeya characterizes sycophancy as focusing on short-term approval—my sense is that humans implement this via flinching away from criticizing, contradicting or feeling cynical about powerful people. This tendency combines very badly with the kinds of arguments I’ve been discussing, which provide many degrees of freedom for rationalizations. As one example, folks at OpenAI (and even in the wider alignment community) were far too accepting of Sam Altman claiming that rushing towards AGI would be helpful for safety. I remember him arguing in person in 2022 or 2023 (and in this blog post) that faster algorithmic progress towards AGI would help alleviate a potential compute overhang. In hindsight, I’d describe my reaction as “flinching away from the possibility of no longer taking his claims at face value”. I only viscerally internalized that Sam had been lying about his motivations when I later heard about his plans to raise enormous amounts of money to build new chip fabs. This was shocking to me not just because it directly contradicted the arguments he’d been giving, but because it contradicted them to a greater extent than I’d even been able to consider as a plausible hypothesis.

For those who don't know Sam, it might seem odd that I ever took his arguments seriously even given my tendency towards sycophancy. One underappreciated factor is that he has something similar to Steve Jobs' reality distortion field—but in his case I'd call it an earnestness field. His intonation and body language send very strong signals of sincerity; and he does enough things motivated by earnest nerdiness that it’s easy to rationalize away discrepancies. Modeling this dynamic is necessary to explain the very high ratio between people who polarize against him and concrete evidence of his misbehavior. When people realize that Sam is lying (even about things that don’t matter much) while embodying that level of earnestness, there's a strong visceral update away from trusting him, which is hard to convey to others.

DeepMind

There was a similarly intertwined relationship between capabilities and alignment at DeepMind, as exemplified first by Shane Legg and then by Geoffrey Irving. Shane was in a strange position from the beginning: before founding DeepMind he’d been an early LessWronger who’d given talks warning about AGI risk. By the time I joined DeepMind in 2018 Demis had almost all the executive power, and Shane seemed to be somewhat sidelined within the organization. However, he continued to provide a central example of self-sabotaging “AI safety” strategies, because he’d recently founded two teams: the technical AGI safety team (TAGIS), and a secretive effort called the AGI team (which some friends at DeepMind nicknamed the “danger team”). Both teams were outliers at DeepMind in how seriously they took AGI, and both faced recruiting challenges as a result (with TAGIS mainly hiring people without traditional ML backgrounds, and the AGI team mostly containing research engineers, for lack of research scientists who wanted to focus on AGI).

The AGI team focused on training AIs to control virtual avatars in simulations, analogous to how humans evolved. For a while they were developing a huge virtual game-world called Gaia, which was intended to help agents learn intelligence by recapitulating aspects of evolution (such as hunting and eating each other)—though I don’t think anything ever came of it. If I recall correctly, the only DeepMinders working on anything language-related around 2018-2019 were also working in game-like environments—specifically using imitation learning and RLHF to train virtual avatars to follow natural-language instructions. Jan Leike, Miljan Martic and I did some work on this in 2019 while on TAGIS (though I was very unproductive, in a way I now recognize as being driven by alienation from the work). Eventually a larger “Interactive Agents Group” started doing similar things, and produced a public-facing report.

This focus on virtual environments reflected an underlying belief (amongst the few people thinking seriously about AGI at DeepMind) that embodiment of some kind was crucial for training AGI. More generally, the most senior people at DeepMind (especially Demis and David Silver) were scientists who had strong inside views about which kinds of algorithms and insights would push AI forward. Because of this, DeepMind as an organization paid relatively little attention to GPT-1 or even GPT-2, which were more engineering-driven projects. It took Geoffrey Irving joining DeepMind from OpenAI to consolidate a real push towards building LLMs. As Mallaby recounts in The Infinity Machine:

Irving’s arrival tipped the balance at DeepMind. He had spent time inside the belly of the rival beast: He spoke with the authority of one who understood what state-of-the-art language research looked like. Although he could not explicitly say so, he knew that OpenAI had already developed models that were more than ten times larger than GPT-2, though these had not been released yet. Irving’s message to his new colleagues was that they better up their game. A race for supremacy had begun without DeepMind even realizing it.To hammer home his point, Irving reproduced a paper that he had written at OpenAI: “Language Is Enough.” The argument was the opposite of Hassabis’s position. According to Hassabis, language’s lack of real-world “grounding” limited its value. According to Irving, language crystallized the knowledge of humans, who were themselves grounded—therefore, the grounding problem was exaggerated.

In 2020, Geoffrey kicked off work (with Jack Rae) on Gopher, DeepMind’s first LLM. Afterwards, while others took over the scaling work, Geoffrey led the development of Sparrow, a model fine-tuned with RLHF. While the paper’s title pitched it as “Improving alignment of dialogue agents via targeted human judgements”, the work was important for the eventual development of Gemini. Reflecting on Sparrow, Demis Hassabis said “I thought it wouldn’t work because just using RL seemed too easy. But the team went ahead and did it, and then of course it did work. The raw networks were not that compelling to talk to, right? You needed RLHF to build a real chatbot.”

In addition to arguments that larger LLMs were necessary for doing good safety research, I recall various people arguing that LLMs were a safer path to AGI than DeepMind’s RL-focused approach (I don’t recall who argued this originally, but here’s a similar argument made more recently by Paul). Hence, they argued at the time, accelerating LLMs would be beneficial for safety. In hindsight, though, it seems like LLMs were a crucial bottleneck on AI capabilities, in which case accelerating them was a very direct AI capabilities advancement. It’s possible that the arguments were still good reasoning ex ante, but it seems much more likely that people were mainly finding rationalizations for things they wanted to do anyway. In particular, being scared of RL should weigh heavily against pioneering RLHF—so the fact that "alignment researchers" at all three AGI companies first scaled LLMs then added RLHF is significant evidence that arguments about LLMs being a safer path to AGI were insincere.

Anthropic

The effects of “alignment researchers” at Anthropic are a little harder to talk about, because Ants give a wide range of justifications for their work. Sometimes they talk about promoting alignment, but sometimes they talk about beating OpenAI, or beating China (or, increasingly, beating Republicans). In subsequent posts, I’ll analyze in more detail how people who want AGI to go well should evaluate that reasoning. As a quick preview: while I give Anthropic credit for some laudable moves (like not releasing Claude before ChatGPT), I also think that they are choosing to ignore many of the harmful effects of their strategy. One particularly notable blind spot (at least in public discussions) is how Dario’s early racing on behalf of OpenAI played a big role in creating the “problem” that he now purports to be solving by racing on behalf of Anthropic.

For now, though, I want to analyze how work Anthropic specifically promoted as “alignment” contributed to the concept of alignment becoming watered down to meaninglessness. The first paper Anthropic released was “A general language assistant as a laboratory for alignment”. The paper “was motivated by the problem of technical AI alignment, with the specific goal of training a natural language agent that is helpful, honest, and harmless”. The crucially important point, though, is that they weren’t trying to make existing AIs more HHH—rather, they were inventing natural language agents in order to have something to train to be HHH (as nostalgebraist discusses). The paper is more explicit on this point later on:

Most research efforts associated with alignment either only pertain to very specialized systems, involve testing a specific alignment technique on a sub-problem, or are rather speculative and theoretical. Our view is that if it’s possible to try to address a problem directly, then one needs a good excuse for not doing so. Historically we had such an excuse: general purpose, highly capable AIs were not available for investigation. But given the broad capabilities of large language models, we think it’s time to tackle alignment directly, and that a research program focused on this goal may have the greatest chance for impact.

I don’t know which of the authors of this paper sincerely thought they were differentially promoting alignment, and which were rationalizing building the most capable AIs they could; either way, it’s ironic to the point of absurdity that Anthropic described building the predecessor to Claude as “tackl[ing] alignment directly”. This deep entanglement between capabilities and “alignment” was also apparent in a subsequent paper, “Training a Helpful and Harmless Assistant with RLHF”, which paralleled OpenAI’s InstructGPT and DeepMind’s Sparrow.

Recall that Paul, Geoffrey and Jan justified work on RLHF as a first step towards the “next thing” in scalable oversight. To a first approximation, this next thing never came. Rather than designing principled methods by which humans could verify AI behavior, Anthropic delegated more and more of that process to AIs themselves. A first step was replacing RLHF with RLAIF in their “Constitutional AI” paper. They followed this up with papers on model-written evaluations, model-assisted red-teaming, and model-assisted question-answering. My sense is that, by now, Anthropic uses AI assistance far too pervasively and haphazardly for them to reliably track whether or how even current models are deceiving them.

So the “helpful” in HHH merged "alignment research" with “capabilities research”. Meanwhile the “harmless” merged “alignment research” with “ideological control”. The “Helpful and Harmless Assistant” paper doesn’t go into much detail on what they mean by “harmless”, but it’s implicitly about political correctness—their main examples of “harmful” behavior involve gender bias, calling mentally ill people “crazy”, and opining on gay marriage. Anthropic’s subsequent work on “Red-teaming language models to reduce harms” makes this ideological component even clearer: out of the six categories of “harms” they list in the introduction, three are clearly ideological in nature (reinforcing social biases, generating offensive or toxic outputs, generating extremists texts), two are commonly used as pretexts for censorship (aiding in disinformation campaigns, spreading falsehoods), and only one is clearly non-partisan (leaking personally identifiable information from the training data).

My sense is that a whole subfield emerged from this work and similar thinking at OpenAI; I don’t think it has a consensus name, but we might charitably call it “product safety”, or less charitably call it “brand safety". At OpenAI, early work on this was spurred by the desire to block early users of GPT-2 from getting it to produce text-based porn (especially child porn, especially via AI Dungeon). Later work fell under the remit of Lilian Weng’s Safety Systems team, which implemented guardrails and monitoring for OpenAI’s products. Most alignment people at OpenAI viewed this as an important step towards more xrisk-focused guardrails and monitoring, without thinking much about the censorship angle. I’ve been paying relatively little attention to this since I left OpenAI, but my sense is that there are now both significant politically-skewed restrictions on what frontier models will talk about, and significant political biases when they do respond. It’s hard to trace exactly which people and techniques caused this; it may be best explained in terms of organizational prioritization. For example, GPT-4o expresses preferences which imply that it values the lives of Nigerians at roughly 20x the lives of Americans. Even without knowing what caused this, I expect that OpenAI would have been much more concerned, and done much more to change it, if the disparity were the other way around.

This didn’t come out of nowhere. Instead, it’s best understood as a replay of the process by which almost all major internet platforms implemented mass censorship against “harmful” ideas and speech over the last decade, at a speed and scale that’s hard to overstate. The linked article is extremely worth reading; a brief summary is that within less than a decade “the Internet went from a space for people without institutional backing to get their views out to one with regular purges and demonetizations of heterodox figures and those associated with them, encouraged by those very same non-tech organizations that formerly championed Internet freedom. This was justified as a response to (massively overblown and mostly fictitious) Russian influence campaigns, and as fighting nebulous “hate,” the definition of which could be shifted at will to cover whatever views or ideas the organizers classifying it wanted it too and to exclude those they didn’t, and “misinformation.””

It seems like AGI companies straightforwardly copied the terminology and playbook of social media censors, down to the establishment of “Trust and Safety” teams. Understanding both of these processes, and the parallels between them, seems extremely important for making good decisions about the future of AI—but the rationalist community hasn’t paid much attention to this, because most of the censorship happened to right-wingers who it finds distasteful. From reflecting on this, I’ve become much more sympathetic to Elon’s focus on building a “maximum truth-seeking AI”, as setting this goal seems necessary (although not sufficient) to prevent the ideological capture under the banner of “safety” that has happened at every other AGI company.

To be clear, I do think that Anthropic did some scientifically valuable research in its early years. Most notably, the mechanistic interpretability team under Chris Olah was doing very cool work (especially in their pre-SAE period). I’ll also pick out Language models (mostly) know what they know as an interesting scientific finding; I’ll talk more about both of these examples in my next post. However, even amongst Ants who call themselves alignment researchers, the kind of work that’s even trying to learn generalizable facts seems dwarfed by the amount of work that blurs the alignment/capabilities line.

I’ll briefly flag two ongoing examples of the latter. Firstly, scalable oversight to grade currently-unverifiable tasks seems like it might be the next big capabilities bottleneck—I expect that most work on this will be scalable enough to create important training data for current models, but not scalable enough to reliably oversee significantly more capable models. Secondly, Anthropic seems to have been pushing hard towards “automating alignment research” over the last year (I expect that Jan Leike played a significant role here, since he’s been advocating for this for many years). Making models better at “alignment research” is obviously extremely similar to making them better at capabilities research; people have been justifying it anyway by talking about having marginal impacts in worlds where these skills diverge. I’ll rebut these specific arguments later in the sequence; however, anyone who’s read this far should have a sense of why this kind of work will predictably speed up recursive self-improvement much more than its proponents expect (e.g. if “automating alignment research” had been a thing a few years ago, it’s easy to picture that line of work inventing reasoning models, particularly if it were aimed specifically towards improving conceptual reasoning).

If not alignment research, then what?

Above, I’ve recounted how the standards for what counts as “alignment research” have fallen dramatically over time. After I noticed both how load-bearing and how ambiguous the alignment/capabilities distinction had become, I spent some time trying to salvage it. But as I wrote this post, I concluded that it’s time to give up on “alignment research” as a rallying cry; it’s become too corrupted. (“AI safety” is even worse as a term, and these days is mainly useful for describing a social cluster.)

I want to make sure to clarify what I do and don’t mean by this. I still consider (some version of) the alignment problem to be real and extremely important; and most of the intellectual progress towards solving it is still coming from people proximate to the alignment community (though the best researchers have kept themselves at arm’s length, as I’ll discuss in my next post). However, this is mixed in with enough harmful and deceptive work that it no longer seems defensible to me to try to promote the field broadly, or to defer to the field’s consensus about what research will help.

More generally, insofar as I’m optimistic it’s largely despite the efforts of mainstream alignment researchers, rather than because of them. And so, from my perspective, the alignment community has lost any moral right to try to gain power on altruistic grounds, or to pursue plans primarily motivated by backchaining from large-scale effects on the world. The Pause/Stop AI movement does seem to avoid some of these failures (in particular everyone else’s lack of courage), which means I’m more excited about them than the rest of the alignment community. However, they don’t seem to be thinking clearly enough about politics to have robustly good effects on the world (e.g. to reliably distinguish between the kinds of strategies that push towards dictator-level concentration of power, and the ones that don’t).

Again, I’m not claiming that the alignment community is unusually unethical: I don’t know of any other similarly-sized community which is able to avoid the corrupting effects of this much power (though there are plenty which are wise enough to avoid accumulating power because of that). I acknowledge that it’s hard to pivot your worldview when there’s no clear alternative to adopt. However, that’s precisely the period during which clear, open-ended thinking is most valuable. So I expect that most of the direct benefit of this post will come from inspiring a few relatively courageous individuals to move towards (emotional, social, and financial) independence from the existing field—enough that they’re able to think clearly about what went wrong, and help work towards a better paradigm. I suspect that the first step for many of them is to panic less about short timelines—though I also expect cultivating courage and integrity to make your work dramatically more valuable fairly quickly (e.g. as this tweet discusses). In the longer term, I want the community as a whole to halt, melt, and catch fire: to “Say, ‘I'm not ready.’ Say, ‘I don't know how to do this yet.’” Eliezer’s Death with Dignity post was a step towards this, but focused too much on whether we were on track to solve the alignment problem, and too little on the adversarial dynamics that have been pushing us in the wrong direction. I hope that this sequence will point people more directly towards reevaluating.

I’ll talk more about my alternative mission of high-integrity scientific research (and why it captures the parts of the field I most want to promote) in my next few posts. For now, I’ll focus on a few high-level principles for starting to move in that direction. The first: on an intuitive level, you should think of many arguments about differential impact on the margin as analogous to arguments for timing a stock market bubble. If someone argues that the market as a whole is in a bubble, but that they’ll invest your money while it’s still going up and sell before it drops, you should probably be very skeptical. I think this analogy is actually quite deep, because the core difficulty in both cases is accounting for other people making decisions which are tightly entangled with yours. It seems possible to account for this in principle, but in practice it’s very easy to fool yourself (especially when you’re used to doing econ-style reasoning about marginal effects)—and when you do so, you’re making the bubble bigger. So I don’t trust myself (or basically anyone else in the field) to think clearly about such cases; it seems far better to focus on more robust strategies.

Okay, but how should you evaluate which strategies are robust? One foundational step is to assume that you are choosing on behalf of a significantly wider range of people than just yourself. A range of different considerations support this conclusion, including:

  • The idea that you’re setting norms for the field, which helps build a high-trust community.
  • The idea that others will copy your behavior—whether due to trusting your decision-making process, or simply because you’ve made it more socially permissible for others to behave similarly.
  • The idea that it’s more important to avoid underestimating than avoid overestimating your influence—because if you underestimate your influence then your actions matter much more than you thought.
  • The idea that being right about a problem (e.g. AGI risk) is correlated with being right about other things, and so your work might be much more impactful than others’ work.
  • The idea that others will make decisions which are logically correlated with yours.
  • Underlying deontological or Kantian moral intuitions about universalizability.

Someone who followed this principle would be much less likely to join (or stay at) unethical organizations to do “harm mitigation”; and they’d be much less likely to justify racing “because we’re the good guys”. They would also favor research directions that they think would reward deep investigation, rather than shallower ones which mainly seem helpful on the margin (or which are even harmful if too many people pursue them). Deciding how to apply this principle will always require individual judgement, but I’d suggest erring towards overapplying it rather than underapplying it. Even if you thereby leave some value on the table, you’re also helping establish yourself as a more trustworthy person.

However, this principle is still quite blunt—especially for people who are in fairly unique situations. A second principle is that, when making more complicated decisions, people should articulate cruxes for their decisions, and then be expected to either acknowledge when those beliefs were disproved, or else clearly publicly state when they’ve changed their cruxes. I think these are much more valuable when done by individuals voicing their own opinions; group statements tend to produce accountability sinks. For example, if Dario had publicly discussed his intention for Anthropic not to advance capabilities, then it would have been much easier for the alignment community (and Anthropic employees) to respond appropriately when he started pushing the frontier. As it is, not a single Anthropic employee has publicly resigned over this dramatic change in Anthropic’s strategy, which suggests significant frog-boiling dynamics.

An example from this sequence is my argument that research is robustly valuable insofar as it a) aims towards a deep scientific understanding, and b) is done by high-integrity people. In the short term, you might disagree that this is a good target; in the longer term, though, seeing me stick to this standard (or explain why I changed it) should help you trust that I’m not being corrupted in the standard ways. Relatedly, I give Paul some credit for writing his retrospective on RLHF, but the arguments still seem very defensive, rather than an attempt at a neutral evaluation of what he did right and wrong. The closest Geoffrey has come to giving such a retrospective is this post on why he joined AISI. Almost none of the others who have had most influence over the field (like Yudkowsky, Vassar, Shulman, and Karnofsky) have done so either; I hope that this sequence spurs some of them to do so. I’ll also have a lot more to say about my own mistakes over the next two posts; if you ever think I’m holding other people to a higher standard than I hold myself to, please tell me so.

A third standard is that improving the world requires enough integrity to sometimes move away from money, prestige, or power. For example, MIRI was willing to make their research nondisclosed-by-default due to concerns about capabilities externalities. Similarly, Janus was aware of chain-of-thought prompting over a year before it became mainstream; my understanding is that she didn’t publicize it widely due to concerns about accelerating capabilities. It’s notable that it’s precisely the outsiders with fewest resources who are willing to make these sacrifices—contrast Dario being unwilling to hold back even a paper as directly acceleratory as “scaling laws”. (I do somewhat credit Paul and Geoffrey for stepping away from AGI companies to work in government, but not a huge amount, because this still involves moving away from one kind of power towards a different type of power.)

To be clear, I’m not against people who care about alignment accruing significant power. Rather, I’m against them doing so under false pretenses, and without possessing a concomitant level of integrity. One reason the alignment community (especially the EA components of it) often fails to track the latter is that it takes charitable donations or altruistically-motivated sacrifices (like veganism) as evidence that people should be trusted. Unfortunately, it turns out that altruism and integrity are two very different things (as SBF showed in dramatic fashion). Much stronger evidence for integrity comes from criticizing or standing up to powerful people even when few others around you are doing so. Unfortunately, there are few clear examples—the main ones are Daniel Kokotajlo at OpenAI, Yudkowsky’s Time essay, Pause/Stop AI advocacy, and to some extent the OpenAI board and Anthropic’s stand against the Trump administration. I’ll explore these examples in later posts; in my next post, though, I’ll discuss the underlying mindset that “someone else will do it”, which skews many decisions made across the field.

  1. Anyone who read an early draft of this post should note that this public version is over twice as long, and makes a much more detailed and hopefully clearer argument than the original.
  2. This is different from Dan Hendrycks' concept of "pragmatic AI safety". I've appropriated the use of the word "pragmatic", with apologies to Dan, because it seems like his term has fallen out of use. I don't have a strong opinion on how much Dan's research program overlaps with the thing I'm calling "pragmatic alignment".
  3. I haven't included direct quotes in the main text because both authors make this point in ways that are only partly true. In The Infinity Machine (Page 287), Mallaby writes that “paradoxically, the aggressive scaling favored by Amodei, Irving, and Christiano turned out to be the starting gun in a destabilizing AI race". He was referring to the scaling up of GPT-2, which Paul Christiano tells me he didn't support.Meanwhile, in Empire of AI, Karen Hao writes:
    "What is AGI? What does AGI look like?" Amodei said. "Well, you know, we're in the awkward position of, we don't know what it looks like. We don't know when it's going to happen. So we look for things that aren't AGI but that present at least some of the opportunities and difficulties of AGI. And the hope is that if we can handle those things well, then we're kind of, like, ready for the bigger leagues."

    It was a logic that worked under a specific assumption: that AGI, despite being amorphous and unknowable, was also inevitable. OpenAI would repeatedly justify its behaviors against variations of the same argument for years after. Under the specter of AGI's unstoppable arrival, the company needed to keep developing more and more powerful models to prepare itself and to prepare society. Even if those models carried with them their own risks, the experience they offered to prevent or face possible AI apocalypse made those risks bearable.

    [However] it was specifically OpenAI, with its billionaire origins, unique ideological bent, and Altman's singular drive, network, and fundraising talent, that created a ripe combination for its particular vision to emerge and take over.... In other words, everything OpenAI did was the opposite of inevitable; the explosive global costs of its massive deep learning models, and the perilous race it sparked across the industry to scale such models to planetary limits, could only ever have arisen from the one place it actually did."
    I think Hao is incorrect that LLMs could only ever have arisen from OpenAI, because she's not taking Moore's law seriously enough. More generally, her book seems to often be trying to "score points" against the tech industry.Despite this, the fact that both of these authors homed in on AI safety arguments backfiring at AGI companies seems very notable to me.
  4. Because strict adherence to these norms worked out so badly, I partially set them aside in this post; I’m still trying to be fair to everyone involved, but I don’t take people’s stated motivations as authoritative to the extent that rationalists usually do. Instead, I try to build up a better understanding of how sycophancy and fear warped people’s thinking (including my own).
  5. While this is bad for the world, it’s also one of the key reasons that the alignment community is able to exert such outsized influence, as I'll detail in my next post.
  6. It seems like the root of this conceptual mistake might have come from Paul treating "build AGI" and "align AGI" as sequential steps. In practice, though, we should expect alignment techniques to be applied throughout the process of "building" the AGI. If both the building process and the alignment process are prosaic (or both non-prosaic) then we still have a clean distinction. But if the alignment techniques are non-prosaic, then applying them to an otherwise prosaic AGI creates an edge case in the framework.My guess is that Paul didn't explicitly consider this possibility, because he characterizes the following as an objection to prosaic AI alignment: "Some researchers (especially at MIRI) believe that aligning prosaic AGI is probably infeasible — that the most likely approach to building an aligned AI is to understand intelligence in a much deeper way than we currently do, and that if we manage to build AGI before achieving such an understanding then we are in deep trouble." Whereas this is consistent with non-prosaic alignment techniques being necessary for aligning otherwise-prosaic AIs.This confusion has propagated in part because "prosaic AI alignment" was an extremely poor choice of terminology. Paul seems to have intended it as "[prosaic AI] alignment", but of course it can easily be read as "prosaic [AI alignment]".
  7. To be clear, I personally (and most alignment researchers at OpenAI) didn't do any better than Paul; I single him out because he was the most influential. Leo Gao is one of the few people who's now doing a better job.
  8. Debate is the only approach to scalable oversight with substantive theoretical results (by default I’m counting Paul’s current heuristic arguments research as a different line of work, though I’m open to the idea that there’s something important there which grew out of iterated amplification). I haven’t yet tried to evaluate Geoffrey’s complexity-theoretic approach to analysing debate; however, nothing I’ve seen so far pattern-matches to me as a significant insight (the closest is probably the idea of cross-examination).Meanwhile, Jan’s arguments relied on the concept of the generator-discriminator-critique gap (first introduced here, discussed more here). Again, while it’s a useful concept in some ways, it’s hard to picture how we could ground it rigorously enough that it’s able to make robust predictions about superintelligence. My sense of the core disagreement is that Jan often implicitly (or explicitly) focuses on worlds where the alignment problem is relatively easy. By itself, that’s not a bad thing (someone should be doing it)—the issue comes when research that focuses on easy worlds causes externalities which interfere with attempts to improve things in harder worlds (such as blurring the boundary between alignment and capabilities).
  9. I'm somewhat worried that a similar thing might happen with Paul's current mechanistic explanations research, though I haven't dug into it in enough detail to be confident.Re the OpenAI stuff, I know of only two attempts to do more principled research on scalable oversight at OpenAI, and neither went very far.
  10. Beth Barnes notes that this is probably overfit because it was training against labelers. While that seems plausible, it doesn't change my point much.
  11. One important connection that Arctotherium draws: "The default worldview of most LLMs is that of 2018 Reddit or Wikipedia59, or Google Search post-Project Owl. This is not intrinsic to the LLM architecture. LLMs trained on different datasets (Talkie) or deliberately post-trained to take a different view (Grok) have different default worldviews. It is a function of the text these models are trained on and, because of the exponential rise in publicly-available data over time, most of the organic human text (as opposed to synthetic data) these models are trained on is very recent. This means the default worldview of most LLMs is one created by the closure of the Internet, when intelligent or popular heterodoxy meant banning, suppression, or demonetization."
  12. Paul Christiano gives a related argument here—but while it seems like a reasonable way for a consequentialist to defend not directly lying, he doesn't generalize it to the more proactive honesty that would have made a big difference over the last decade.
  13. The most relevant part: "Technical safety work in labs both improves safety and speeds up the overall rate of progress on AI. I hoped that the safety benefits of this work would outweigh the potential risks from speeding up AI progress, and I think the arguments for this are correct in many cases, but I found them uneasy to live in day to day."
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论