The Lens That Conceals Its Flaws: Nonsense at the Heart of the Sequences

Sourced from my Substack and adapted for LessWrong. Epistemic status: confident.

No online community has shaped how we think about AI as much as LessWrong. The community Eliezer Yudkowsky founded now runs through the whole field. Longtime LessWrong posters are building the future at frontier AI labs, charting AI capabilities, researching alignment, and prophesying imminent doom if we don't throttle progress.

LessWrong conceives of itself as a college-like intellectual community with a distinct educational mission. The founding text of that mission is the Sequences, Yudkowsky’s long series of essays on rationality. The Sequences are so important to LessWrong that every new poster is advised by moderation bots to read Highlights from the Sequences to learn the ropes.

The Sequences have arguably done rationality a service by popularizing concepts from philosophy and psychology that are genuinely helpful for being more rational. What they do much less well is practice what they preach. Rationalism as applied in the Sequences is sometimes fundamentally absurd, and it seems quite plausible that the confused thinking at its foundations has led to less-than-helpful perspectives on the increasingly mainstream, high-stakes questions of AI risk.

A closer look at the Rationalist lens

In the first essay of Highlights from the Sequences, The Lens That Sees Its Flaws, Yudkowsky makes a big promise: I will give you a way to inspect and debug the mental machinery that produces your beliefs. Towards the end of the essay, he gives us a look at how he applies his method:

Once upon a time, I went to EFNet’s #philosophy chatroom to ask, “Do you believe a nuclear war will occur in the next 20 years? If no, why not?” One person who answered the question said he didn’t expect a nuclear war for 100 years, because “All of the players involved in decisions regarding nuclear war are not interested right now.” “But why extend that out for 100 years?” I asked. “Pure hope,” was his reply.Reflecting on this whole thought process, we can see why the thought of nuclear war makes the person unhappy, and we can see how his brain therefore rejects the belief. But if you imagine a billion worlds—Everett branches, or Tegmark duplicates—this thought process will not systematically correlate optimists to branches in which no nuclear war occurs.To ask which beliefs make you happy is to turn inward, not outward—it tells you something about yourself, but it is not evidence entangled with the environment. I have nothing against happiness, but it should follow from your picture of the world, rather than tampering with the mental paintbrushes.

This passage contains some truth: wishful thinking isn’t evidence that the wished-for thing is true or will happen. The issue is how Yudkowsky gets there, and whether the method he’s demonstrating is one we should trust. This is not a minor issue, but the heart of the topic at hand. If rationality is anything at all, it is the practice of reliable methods of reasoning. The Sequences are intended to teach it, and they must teach by example if they are to deliver on their central promise.

Using esoteric nonsense to “prove” what we already know

Slow down, and really look at what Yudkowsky’s argument says. It asks us to imagine a billion worlds and “see” that hope doesn’t “correlate optimists” to worlds where nuclear war is avoided. There are two problems with this:

  1. No one, including Yudkowsky, has ever actually done it in any meaningful sense. Any human who claims they can take a mental sample of a billion possible worlds and calculate statistical correlations on that sample is, charitably, bullshitting for entertainment purposes. Rationality is a mental activity where we make progress by doing things with our minds, so the things have to be possible to do.
  2. Even if you suspend disbelief and pretend to do Yudkowsky’s thought experiment anyway, the result has no value, because it only tells you what you already believe. We already know that hope doesn’t do anything in a context like this. If a random person hoping for peace somehow did make peace more likely, the thought experiment would mislead us, because it doesn’t give us access to anything outside our own heads.

If contemplation of possible worlds doesn’t tell us why wishful thinking fails to make things happen, then how do we really know that? Because, if hoping did make things happen, we could do things by lying around hoping instead of moving our bodies. Every person on Earth has tried this experiment many times, and we all know it doesn’t work. That is the actual ground of the conclusion: direct experience of how the world responds to our intentions, not introspection into a mental picture of the multiverse.

Furthermore, because the thought experiment is purely introspective, it silently fails to reveal the nuances that only a real-world investigation can provide. Nuclear war is one of the clearest cases in social science where beliefs can shape outcomes. According to Thomas Schelling’s analysis of deterrence, if each side in a potential nuclear confrontation expects the other to strike first, the incentive to strike first grows, and the expectation of war becomes a cause of war. Robert Merton coined the phrase “self-fulfilling prophecy” for dynamics of this kind.

The mental state of one private person generally doesn’t move the odds of nuclear war, but the real-world effects of hope at large are revealed only by engagement with the outside world, not just the contents of one’s own mind.

The Sequences sometimes deliver mystification

In 2008, Deena Weisberg and colleagues found that adding irrelevant neuroscience to explanations of psychological phenomena made non-experts rate them as more satisfying, and the boost was largest for the bad explanations. Irrelevant technical content is a form of mystification that can lend false authority to whatever it’s attached to. Everett branches in an essay about basic epistemology do the same thing.

Worse, they conceal the empty circularity of the argument behind a showy conceit of near-divine insight. If Yudkowsky had asked us to remember ordinary situations, he would have gotten us straight to the actual grounding of our belief. Asking us to imagine ordinary situations would be less effective, but at least it would be clear that we were just reminding ourselves of our own beliefs and learning nothing. Everett branches are the worst possible choice. On the many-worlds interpretation of quantum mechanics, they are physically real parallel universes, and imagining a billion of them seems like a godlike act, akin to directly perceiving or even creating reality at the ultimate scale. This is not only needlessly obfuscated, but carries considerable risk of grandiose self-deception.

Mystification is an ancient human trick that some may see as harmless fun, or even useful. A great illustration of its strange effects comes from the 2001 B. R. Myers essay A Reader’s Manifesto, which criticized the trend towards baroque imagery and prose in literary fiction. Along the way, Myers quotes an early biography of Edward Pococke, a seventeenth-century English parson who was also one of the foremost Arabic scholars of his age, and who insisted on preaching so that his rural parishioners could understand him:

But from this very exemplary caution not to amuse his hearers (contrary to the common method then in vogue) with what they could not understand, some of them took occasion to entertain very contemptible thoughts of his learning ... So that one of his Oxford friends, as he traveled through Childrey, inquiring for his diversion of some of the people, Who was their minister, and how they liked him? received this answer: “Our parson is one Mr. Pococke, a plain honest man. But Master,” said they, “he is no Latiner.”

One of the most learned men in England was dismissed by his own congregation because he spoke plainly. Weisberg’s study and Myers’s anecdote both suggest that people can be made to prefer mystification if it is delivered in a familiar, socially prestigious argot. In Pococke’s England, the prestige argot was Latin quotation in sermons. Today, the vocabulary of science is one of the most popular means of mystification.

Mystification is disempowerment

Physics education research shows that learning challenging material can be insidiously disrupted by subtle failures to align the method of student engagement with the precisely defined learning goal. Derek Muller found that students rated clear explanations that seemed to match their intuitions highly but learned little from them, while explanations that confronted their misconceptions felt confusing yet produced real gains on tests. Louis Deslauriers and colleagues found the same pattern at Harvard in 2019: students in active-learning classes learned more but felt they had learned less than students in polished lectures.

This creates a dilemma for teachers. They’re often evaluated based on student reports, but student reports aren’t always aligned with actual learning. Real learning often feels like struggle, which can make students feel worse temporarily. If you want students to report understanding, that’s more easily induced by combining an entertaining format with subtle signals of the status the student will attain from appearing to have advanced technical knowledge. Intentionally or unintentionally, Yudkowsky’s passage is a supernormal stimulus optimized for precisely this experience of rewarding but fake learning.

The ultimate aim of the Rationalist project is to prevent us from inadvertently rewarding AI for thinking and doing the wrong things, leading to an eventual annihilatory turn against humanity. When we juxtapose this noble goal with the disguised human disempowerment its texts sometimes actually deliver — almost certainly inadvertently — this is an irony of grievous proportions.

Education in rationality is empowering, and should be a public commitment

The point of rationality, if we actually care about getting people to practice it, is that much of it is not rocket science. Anyone can do it and make their life better, just like anyone can learn to make a sandwich or get a yellow belt in karate. There are much higher levels of skill that aren’t for everyone, but there’s a basic citizenship level, mere rationality perhaps, that makes the world a better place if everyone does it. That’s the level we should be popularizing. And when we hire for that job, we should evaluate candidates just like we’d hire a math teacher: they must be able to do the work.

I have some thoughts about how we can find good teachers, but I’ll leave drawing them out to another day, or to the professionals. For now, like Edison on an average day in the lab, I’ll have to be content with having found yet another way to rationality that doesn’t work.

  1. As one LessWrong moderation message states: “LessWrong has fairly specific standards, and your first LessWrong post is sort of like the application to a college.”
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论