Why I believe current AI models can really think and understand - Part 1 of 2

This one’s long. To make the complete argument here, I had to introduce a lot of ideas from the philosophy and science of mind. I figured if people would read a novella about data centers and water, they’d read one about what might be the most important philosophical question right now. If I had a few hours with a smart friend to get across what I actually think about AI and philosophy of mind, this is the very very best I could do.

Introduction

Okay, first of all, I want to be clear that when I say AI models can “think” and “understand” I’m not claiming that they have minds like ours. People use “consciousness” to mean two related but different things, which the philosopher Ned Block calls “phenomenal consciousness” and “access consciousness.” Phenomenal consciousness is the first-person experience of what it’s like to be you. We usually think of this kind of consciousness as something more than information processing. There’s some unique experience that can’t be translated into a pure third-person account of your brain states. Access consciousness is information in your brain being available to lots of parts of your mind at once. This one IS fully describable in the third-person language of information flowing through a physical system. An AI could have access consciousness (sharing information across lots of its internal processes at once) without having phenomenal consciousness. There wouldn’t be “anything it’s like to be” it. Being that AI would “feel” just like being dead. I won’t assume or argue that AI has phenomenal consciousness.

I’ll make two claims, one in each post:

  1. Our thinking and understanding mostly don’t depend on consciousness, and where they do, the aspects of (access) consciousness they depend on could in principle be replicated in AI models. (This post)
  2. Current AI models seem to have the qualities that are necessary and sufficient for real thinking and understanding. (Part 2)

These claims are controversial, but I’m pretty convinced by the arguments for each. My argument won’t rest on any one controversial belief about the nature of thinking or AI. Instead, I want to show that most leading theories about the mind seem to strongly imply both claims, or at least give them a high enough chance of being true that we need to take them seriously.

Importantly, I won’t argue that AI can do all the thinking we can do, just that what AI does at this point fits my definition of real thinking and understanding. I can’t think in the way Einstein could, but that doesn’t mean I can’t think.

You might think this is all just arguing over semantics, and it doesn’t matter what words we use for AI, because it can do what it can do either way. I think this is wrong, because most people who say AI can’t “really think” go on to infer a lot from that. They conclude it won’t be able to “discover anything new,” or that it’ll be permanently limited in how much it can change the world. If it’s “just linear algebra” it will remain a fun toy. If it can’t think, it certainly can’t “plan,” so it might not be especially dangerous. The most fundamental inference seems to be that AI will never do well when it has to carry a line of reasoning through to a case it hasn’t seen before. These are the situations where only the most general, subtle patterns from past experience can help, and only when a careful conscious person deliberately thinks them through. This, above all, is what a lot of people mean by “thinking,” as opposed to regurgitating information. Getting this question right matters a lot for making sense of where AI is headed.

I’m also worried that as debates about AI have heated up, a lot of people have adopted a rule of thumb that any talk of AI “thinking” is effectively assuming that there’s a ghost in the machine, like assuming a radio has a little band inside it playing the music. This seems like a wild overreaction to me, and one that wouldn’t make sense outside our current cultural moment with AI. I’m pretty convinced that if I took a current model like Astra or Opus 5.5 back to 1960 and let people play around with it, the reaction from serious people would be “Wow! You’ve created machines that can actually think! They understand the words they’re using!” Alan Turing predicted in 1950 that by the end of the century, “the use of words and general educated opinion will have altered so much that one will be able to speak of machines thinking without expecting to be contradicted.” Today, people are much more wary of saying this, partly because they think “anthropomorphizing” AI buys into the hype or is a sign that you’ve developed psychosis. I think that’s an understandable (initial) reaction, but it’s starting to cause people to make serious mistakes. Part of my argument will be that the hard rule against “anthropomorphizing” AI is itself philosophically confused, and potentially anti-scientific, because it often implicitly depends on and enforces pre-scientific folk theories about what’s special about human minds.

It’s bad and dangerous to use anthropomorphic language for AI when it leads people to believe that AI has qualities it doesn’t. If someone comes to feel an AI really cares about them, or that it has a human-like mind, that seems pretty destabilizing, and it’s understandable that people worry about it. But the worry can go too far.

Before machines existed, humans and some animals were the only things that lifted objects. Once we built machines to help, it wasn’t bad to use anthropomorphic language and say a forklift “lifted” something, even though lifting used to be something only conscious, biological animals did. The word “computer” used to be a human job title, but no one today would say it’s anthropomorphizing a laptop to say it “computes.” These are obvious extreme examples. Where’s the line between clear cases like these and cases where it becomes bad to use anthropomorphic language? The obvious answer to me is that the line is where this language implies AI has a lot of abilities it doesn’t have. But the same line also shows where it’s bad NOT to use anthropomorphic language for AI, which is when avoiding it implies that AI lacks abilities it actually has. Ultimately, what makes something human-like is a specific combination of lots of simpler properties that aren’t human-like on their own. Anthropomorphic language is useful because it helps us simplify, understand, and predict the behavior of our fellow incredibly complex systems we call humans. It doesn’t pick out some fundamental unifying aspect of the properties of our minds that walls them all off from anything else in the universe having them.

So deciding when to anthropomorphize AI depends on what abilities it actually has, but also on what abilities our own minds have, and I think most people’s picture of how their own minds work is deeply mistaken. I’ve called this picture “folk-Cartesianism” in past posts. This view is that a lot of aspects of the mind come in a single inseparable package, held together (and walled off from machines) by the fact that all of them require a conscious, first-person perspective observing and managing them. This view imagines the mind as an internal movie theater, with a little version of you inside watching the movie of your life. The philosopher Daniel Dennett called this idea the “Cartesian theater” after Descartes, who often wrote about the mind as an unchanging observer (the “I” in “I think therefore I am”) watching its own experiences and thoughts directly.

The folk theory of mind says that when you “think,” you’re consciously observing individual concepts and moving them around the screen in your mind. You react to them and decide the direction of your thought. A key part of thinking (in this view) is powerful, direct introspection, where you turn your mind’s eye back on your own thought process and observe it directly. Your thoughts and inner experience are clear and directly accessible to you. You can observe them as if they were all written out on a piece of paper you were reading.

The folk theory says that if something can’t do this (if it can’t look at its inner mental screen and see its own thought process laid out clearly), then its “thinking” is basically only a domino chain running from premise to conclusion, like a rote computer program chugging along with no idea what it’s doing. That can be useful, but isn’t thought.

I believe that this picture is a big reason so many people object to the idea that AI “thinks.” If you hold it, “AI can think” smuggles in two drastically wrong (according to the critics) ideas at once. The first is that AI has first-person conscious experience. The second (which follows from the first) is that whatever AI does that looks like thinking is just as good as real human thought, when without that inner glow of consciousness it would have to be missing most of them.

In this post, I’ll argue that this folk theory is deeply wrong, for a few overlapping reasons. In Part 2, I’ll go through what current AI can do. I’ll argue that current AI does enough of what we mean by thinking and understanding that the burden should now be on anyone who says it doesn’t think to say what’s still missing. And after this post, I hope you’ll agree that the list of what’s missing doesn’t include a Cartesian theater (a conscious observer inside the AI with direct, easy access to its own thought processes and control over them), because I don’t think humans have one either. I’ll knock down human thought a bit and raise up AI thought a bit, to the point that they might come close to meeting in the middle.

To be clear, I don’t mean to say here that AI can think all the thoughts humans can, for the same reason that when I say I can think, I don’t mean that I can do everything that any human mind can do. But a core part of how I think about the world is that most of our beliefs about our own minds and thought are useful but wildly imperfect folk theories and heuristics that get at reality about as much as saying “the rock wants to move toward the ground.” I think it’s a gigantic mistake to assume we have total access to what’s going on in our minds, and to let folk beliefs about our minds’ specialness override any debate about them.

How the human mind is different from our conception of it

Some examples of simple thought and understanding that don’t require consciousness

I’ll start with some simple, first-person tests you can run on your own thinking, to show that it’s not entirely like an inner theater. If your conscious mind does all your thinking and has full access to it, it should be pretty easy to watch it happening.

First, think of a city, any city at all.

Do you have one? Why did you pick that one? Do you know?

When I do this, I don’t consciously search through a mental list of every city I know and pick one. Instead, a city just appears, or several appear and I pick one of them. This seems like a pretty clear sign that some part of your brain understood what “city” meant well enough to search your inner catalog of examples without your consciousness overseeing it.

Now we’ll do another, and this time you should pay really close attention to the process in your head.

Pick an ice cream flavor.

Candidates appear, but again it’s not clear where they’re coming from.

Now finish this sentence: “The early bird catches the...”

I don’t think you had any mental control over whether you thought “worm” there. It appeared before you made any conscious decision, and I think it would be very hard to decide not to think of “worm.”

Think about the last time a word was on the tip of your tongue. You might be trying to remember an actor’s name, and can recall their face, the movies they were in, even the letter their name starts with. Then the answer suddenly appears hours later while you’re doing the dishes. Some mental process was humming away without your conscious mind knowing, and was suddenly able to produce a result and deliver it to your conscious mind.

All of these are pretty obvious evidence that retrieving information (which itself depends on understanding a huge amount of context about what words mean and how they relate to each other) can happen without our conscious minds observing it or making any decisions about it.

Folk theories of the mind often imagine language as playing a pretty secondary role to thought. When you speak, the folk theory imagines that you have some clear, pre-linguistic picture in your head of exactly what you want to say, and choose the words that will best communicate that picture to another mind. The full idea is clear before you speak, and language is just a useful bridge between minds, rather than something that our mental pictures depend on.

I’ll challenge this folk idea of language in a lot of ways in these posts, but one immediate problem with it is that the words we say often require a shocking amount of specific reasoning and understanding that doesn’t exist in our conscious minds before we say them. To see this, you can choose a topic you know a lot about, set a timer for 3 minutes, and then start speaking about it as much as you can. As you talk, try to pay really close attention to what’s happening in your conscious mind, and what decisions you’re actually making before the words come out of your mouth. You might be surprised to find that you don’t really know how a sentence will end when you start it, and there’s very little conscious planning happening at all. And yet you speak coherently, and can even learn from the things you hear yourself say. E. M. Forster tells a story about an old lady who asked, “How can I tell what I think till I see what I say?” This seems to imply that a lot of the “thought” required to plan and produce a complex paragraph happens outside your conscious awareness. And because producing it takes a deep understanding of the words and how they fit together into coherent English, as well as the concepts and how to best explain them via carefully structured sentences, a lot of that “understanding” is happening outside your conscious experience too. The fact that this works is one reason why writing can often be useful for learning. Writing can reveal full thoughts we’ve had that we might not have been consciously aware of.

One famous example of complex thought happening outside consciousness comes from the mathematician Henri Poincaré. He’d spent about 2 weeks working hard on a new kind of function. He took a break to go on a geology trip and stopped thinking about the problem. He says:

At the moment when I put my foot on the step the idea came to me, without anything in my former thoughts seeming to have paved the way for it, that the transformations I had used to define the Fuchsian functions were identical with those of non-Euclidean geometry. I did not verify the idea; I should not have had time, as, upon taking my seat in the omnibus, I went on with a conversation already commenced, but I felt a perfect certainty.

When he got back, he checked it, and it was correct! He concluded that sudden flashes of insight like these are “a manifest sign of long, unconscious prior work.”

In all of these examples, your conscious mind only gets access to the output of thinking. It doesn’t observe the processes that produced it. Scientists disagree about how far this goes. The neuropsychologist Karl Lashley had an extreme take in the 1950s, when he wrote “No activity of mind is ever conscious.” The psychologist George Miller later said, more mildly, that it’s “the result of thinking, not the process of thinking” that shows up in consciousness on its own.

So the folk idea that all thought and understanding require consciousness seems wrong. The question is what role consciousness plays in higher level thought and understanding.

A lot of your understanding happens outside of consciousness

So far my examples have mostly been loose associations (like thinking of a city) or drawing on things you already know (like talking about a topic). Maybe a lot of understanding and thought is by definition more complex than this, and requires conscious experience.

But we also understand much more complex things that we’re almost never conscious of, and never consciously learned. For example, English speakers put adjectives in a very consistent order. “A big red balloon” sounds normal, but something sounds strange about “A red big balloon.” The pattern adjectives follow is roughly opinion, then size, age, shape, color, origin, material, and purpose. I could describe Washington DC as a pleasant small old square gray city, and that sounds fine, but saying it’s a gray square old small pleasant city sounds strange and unnatural. You understand and use this pattern every time you speak, but you might never have been conscious of it, even when you “learned” it from years of hearing English. I would say I understand how to use English well, and this pattern is a key part of using English well, yet I went a long time not being conscious of this specific rule.

You also understand which words ought to follow others, well enough that your brain reacts to a strange word in a sentence within a fraction of a second, without you consciously deciding to check whether it makes sense. Neuroscientists have recorded people’s brain activity as they read sentences one word at a time. Some sentences ended with words that didn’t fit, like “He spread the warm bread with socks.” Strange words like “socks” trigger a distinct electrical response in the brain that peaks about 400 milliseconds after the word appears. Later experiments found that words that make complete sense also produce this same response, and the electrical response gets bigger the less expected the word is. So as we read, our brains seem to be constantly tracking how predictable each next word is. One explanation is that our brains use context to anticipate likely words. This seems like a core part of how we understand what we read. It’s unconscious (we’re not consciously weighing every word that could come next), and it looks surprisingly similar to how modern AI “predicts the next word.” This also looks to me like a deep understanding of how lots of possible words relate to the situation at hand, all happening unconsciously.

Some researchers wanted to see if there were any signs of deeper similarities between how our brains predict words and how AI does. AI models also assign a likelihood to every possible next word (that’s how they write), so we can see when a word shows up that the model considers very unlikely, like spreading bread with “socks.” Several research groups have found that these likelihoods inside language models reliably predict how strongly people’s brains react to more or less expected words. And the models that are best at predicting the next word tend to be the best at predicting the brain activity of people reading. So AI seems to successfully mimic at least this part of our unconscious understanding.

Our introspection and memory of our conscious thought are fallible

So we’ve seen there are types of thought and understanding that don’t occur in our conscious minds, which implies they don’t require consciousness. But the folk picture also runs into trouble with the thinking that does happen in our conscious minds. Our very recent mental past is surprisingly opaque to us. Even our honest explanations of our own recent conscious decisions can often be false.

The most extreme example of our beliefs about our own minds being false comes from split-brain patients, who have had the main bundle of fibers connecting the two halves of their brain cut to treat severe epilepsy. Researchers can show these patients a picture that only one half of their brain sees. In most people, only the left half of the brain can generate speech. In a famous study, the neuroscientist Michael Gazzaniga showed a picture of a chicken claw to the left side of a patient’s brain, and a picture of a snowy scene to the right side. The patient was then asked to pick related pictures from a set. The patient’s right hand (connected to the left side of his brain, which had seen the claw) chose a chicken, and his left hand chose a snow shovel. When he was asked why (remember, this is the left side of his brain talking), he gave an answer he honestly believed. He explained that the chicken obviously went with the claw, and “you need a shovel to clean out a chicken shed.”

So the speaking part of him had no idea why his other hand picked the shovel, but it came up with a compelling story, after the fact, for why he’d consciously chosen it. Gazzaniga came to believe that we all have a system like this in our brain, which he calls the “interpreter.” Its job is to come up with after-the-fact justifications for what we do, which we experience as honest beliefs, and to weave them into a coherent story about us, whether or not it has access to the real reasons behind our decisions. If this is right, the reasons behind even our very recent conscious decisions are more hidden from us than they seem, and consciousness plays less of a role in our higher-level decisions than we’d expect.

This hints at a broader claim I want to make, that introspection is just another mental process. It takes in uncertain information, processes it, and gives an output that’s as reliable or unreliable as our beliefs about the outside world. Ultimately, our thoughts are the result of activity in our brain. There isn’t some ethereal realm we exist in where we can observe our thoughts directly. Our beliefs about our own inner lives are fallible stories about the past, made more coherent by the (often self-serving) narratives our brains offer up, which we honestly come to believe. The idea that introspection is some transcendent act of turning the mind’s eye back on itself and seeing its own processes completely clearly is pretty suspicious.

We’ve seen similar evidence for something like an “interpreter” in people with intact brains too. In a 2005 study, people were shown pairs of photos of women’s faces and asked which one they found more attractive. In some trials, the researchers used sleight of hand to swap the photos, so people were handed the face they’d rejected and asked why they’d chosen it. The swaps were only caught in the moment 13% of the time, and only a quarter of the swaps were noticed at any point in the process. Instead, people very confidently and honestly explained why they had “chosen” this face, often indicating features that only the swapped face had. For example, one man said he had chosen the “more attractive woman” because she was wearing earrings, but the woman he had actually chosen wasn’t wearing any. The woman he had been handed after and was looking at was wearing them. 84% of the people who didn’t notice a swap said they thought they would notice one.

So people can honestly come to believe fake narratives their brains offer up about very recent conscious decisions. Dennett believed that this process describes our whole minds. He didn’t think there’s actually a single place in the brain where everything comes together for a viewer to see. He believed the self is basically just the broad story all these individual narratives and justifications add up to, which he called a “center of narrative gravity.” We tell stories about ourselves, but “for the most part we don’t spin them; they spin us.”

This view of introspection makes sense if you think about what actually happens when we introspect. Whenever we think at all, we’re feeding information about our own past behavior back into our cognitive processes.

The output of mental activity can then be used by other mental activity.

All of our mental activity is fallible, and takes in uncertain information about what’s happened before. Introspection specifically takes in past (fallible) evidence we have about our own behavior and computes a (fallible) output.

What doesn’t seem possible is a folk theory of introspection that imagines us as floating observers with a god’s-eye view of everything happening in our own minds:

Ultimately what’s available to the mind is the results of past computations and mental processes. Some of those could involve reporting back to the mind a lot of the individual steps taken. Introspection in my view is like examining any other aspect of the world. This more natural kind of introspection seems easily available to AI, which could also just collect information about what its past processes did and act on it. It doesn’t need a floating eye in the middle of its “experience” looking back at its own processes.

Some people point out that AI models can only move information forward through their layers, and take this as a sign that they can’t do what we do when we introspect. To me, this sounds kind of like saying that human minds can’t introspect because all the processes inside are moving forward in time. I think our folk theory often gives us the mistaken idea that we can directly peer into the recent past of our own mind and examine what’s there. What’s actually happening is that we’re using current mental processes to retrieve information about past ones. AI models can do the same thing, retrieving the results of earlier processes while answering a question (and we know they do this regularly). Once you think about introspection this way, it’s a little hard to see why current AI models couldn’t do something similar.

So it’s looking bad for the ‘little version of you in a theater’ view of the mind. The process that produces your thoughts (which takes understanding) is hidden from your conscious mind, and the process that explains your thoughts to you might be a separate “interpreter” without reliable access to the first one.

What about slow, careful thinking?

Of course, most of my examples so far are quick, automatic kinds of thinking. Slow careful reasoning, like working through a philosophy argument or doing math, looks like a much stronger candidate for where the conscious version of you takes over.

Even here, lots of individual steps happen unconsciously. If you ask me what 7 times 8 is, 56 just appears in my head. Maybe slow thinking works by summoning these opaque individual answers into our attention, and then examining them in the light of consciousness to build a deeper understanding.

Slower reasoning is not our default. It takes more energy and often feels unpleasant (this was the big theme of my last post on EA). In the famous simple problem:

“A bat and a ball cost $1.10 together, and the bat costs a dollar more than the ball. How much does the ball cost?”

The math looks so simple that most people don’t even think to slow down, and only switch into careful reasoning after they hear the ball isn’t $0.10 and get confused.

What actually determines when you enter this slow thinking? Is it a conscious decision? You can definitely choose to think more slowly and deliberately about a problem, but often the switch happens unconsciously. We don’t start reasoning slowly until something “snaps us awake,” like coming off autopilot. A lot of the time, it’s not clear to my conscious mind what aspect of me causes that slower reasoning to kick off.

But once we’re in slow reasoning mode, this does seem like where consciousness matters most for our thinking. It’s hard to picture a machine doing our most important kind of slow, deliberate thinking (pulling in information from lots of contexts at once and carrying very general patterns into new areas) without something like consciousness. So I now need to say what consciousness is (or might be) and what role it seems to play in our most careful thinking.

What phenomenal consciousness is

I’ll rely on Ned Block’s distinction from the introduction, between phenomenal consciousness (the first-person, what-it’s-like experience Thomas Nagel writes about in “What Is It Like to Be a Bat?”) and access consciousness (information being broadly available to lots of processes in your brain).

I am personally pretty convinced that phenomenal consciousness is not required for thinking or understanding. Access consciousness is required, but I think the way it works in humans implies current AIs probably have something close enough to it to allow for high-level thought.

Most of this section will be on phenomenal consciousness, and the next will be on access consciousness.

I think most people see phenomenal consciousness as the clear ultimate barrier between human minds and machines. No matter how advanced machines get, they will never ascend above information processing and have the first-person subjective experience we have of the world. For example, a machine could look at this square, and detect the wavelength of the light it’s giving off:

But it could never have the first-person experience of the redness of the red.

When you look at this square, I could write a third-person description of what’s happening in your brain. I could record the activity of every neuron, the wavelength of light, the associations you’re drawing between the color and past information your brain’s received about apples or sunsets. But no amount of third-person information about your brain, even the position of every particle and every electrical signal, could ever give me the first-person experience of being you seeing the red, the redness of the red over and above the information carried by the light.

This is what the philosopher David Chalmers called the “hard problem” of consciousness, the strange fact that consciousness alone seems to be inaccessible to third-person observation. Nothing else in the universe works like this. Everything else exists somewhere in space and time, follows regular physical laws we can write down as math equations, and can be measured by anyone with the right instruments. More complex “things in the world” like the national debt can ultimately be reduced to the behavior or expected behavior of systems made up entirely of things that have these basic qualities. First-person conscious experience seems to break all these rules. People could dig around in your brain as much as they wanted and find the neurons that respond to red, but it doesn’t seem like they’d ever find the redness.

This seems pretty hard to deny. Unfortunately, another fact about the world also seems really hard to deny, and it contradicts the first one. Physics seems to be “causally closed,” meaning the only things that affect physical objects are other physical things described by the laws of physics.

One of the main reasons to believe this is the conservation of energy, which says that when energy leaves one object, it always goes to other objects. When one type of energy disappears from an object without going to another object, it’s just turning into another type of energy. For example, when a rock falls, its gravitational potential energy goes down, and its kinetic energy (the energy it has as a result of moving) goes up. If energy is never created or destroyed, then no energy can be added to the world from some mysterious realm beyond physics. If nonphysical things were affecting the physical world, objects would gain or lose energy or momentum out of nowhere, breaking these laws. If a ghost that exists outside physical reality picked up a rock, the rock would gain energy that came from nowhere. Classical physics couldn’t work if violations like this were happening all the time, so physicists came to assume they weren’t happening at all.

We now know that classical physics was way too simplistic, and its version of these laws IS regularly violated. Mass and energy aren’t conserved separately. Mass can turn into energy, and energy into mass. Physicists now believe in the conservation of mass-energy. You can turn about a gram of mass, roughly the weight of a dollar bill, into the energy released in the Trinity test (the first atomic bomb), but you can’t turn it into the amount of energy in the sun. There are still strict rules for these conversions, set by E = mc2 and the broader conservation law.

On the standard textbook reading of quantum mechanics, there’s also fundamental randomness in the world that isn’t the result of our not having enough information. If you rewound the universe, started it from the same initial conditions, and played it back, different things could happen. This is different from the everyday probabilities we deal with. When I roll dice, they’re so completely governed by classical physics that if you knew their exact position and velocity, their distance from the table, and exactly how they’d bounce off it, you could predict with perfect certainty what numbers they’d land on. If you rewound the roll and rolled them in exactly the same way, they’d land on the same numbers every time (unless quantum effects showed up at the everyday scale, which is extremely unlikely but not impossible). Our everyday uncertainty just comes from missing information. Quantum mechanics in comparison appears to be truly probabilistic.

Some people mistakenly think that quantum randomness opens the door for nonphysical entities to affect the world. The idea is that something like consciousness, existing beyond the laws of physics, could take an event with a 60% chance of happening and “choose” that it happens. From the outside, this wouldn’t look like a violation of physics. But this deeply misunderstands what quantum probabilities are. For the rest of quantum mechanics to work, a 60% quantum probability has to reflect the true odds of the event. The odds themselves are part of physical law. An outside entity meddling with them would basically be changing the 60% to 100%, which violates physical law in the same way as the ghost lifting the rock. Among other problems, the ability to change these probabilities would mean we could use entangled particles to send messages faster than the speed of light, breaking another core rule of modern physics. So quantum mechanics is actually just as walled off from supernatural effects as classical physics was.

Physics is (I’m biased) our very most successful field of science. It has broken down the nature of ultimate reality and successfully predicted the behavior of distant stars and black holes. Giving up causal closure would break a lot of what makes it work.

So, all that to say, physics looks really really causally closed to me. That leaves us with a problem, because conscious experience doesn’t seem to behave like anything physical. All this means that of these three obvious-seeming statements, one has to be false:

  1. Physics is causally closed. Everything that happens in the physical world is fully caused by other physical things (or, where there’s quantum randomness, has its odds set by them).
  2. Phenomenal conscious experience isn’t reducible to physics.
  3. Phenomenal conscious experience affects what we do, including what we say about it.

Which one you cut matters a lot for how consciousness affects thought, and so for whether machines can think. I’d say my main problem with human existence itself is that cutting any of these looks like complete garbage to me.

The decision to cut 3 is called “epiphenomenalism.” This is the idea that consciousness is not reducible to physics, but physics is causally closed, so consciousness doesn’t actually have any effect on the physical world. Our conscious experience is like a passive observer in a theater helplessly watching our lives unfold.

This can sound intuitive at first, but I think it looks completely crazy the more you consider it.

This view implies that whenever you say “I’m not just processing information, I actually experience the redness of red,” your experience of the redness of red played no part in making you say it. Your physical brain, your speech, and your writing (and anyone else’s) have never been affected at all by the fact that you have this special experience. It’s just a pure coincidence that our physical bodies and brains, which are basically machines without this phenomenal experience on their own, happened to start writing about this thing that has never once affected them but definitely does exist beyond the physical world. You may have heard of “philosophical zombies” (or p-zombies), an idea that fits naturally with this view. A p-zombie is physically identical to a normal person but doesn’t have phenomenal consciousness. There isn’t anything it’s like to be them. An epiphenomenalist believes an entire world could exist exactly like ours, with everyone doing everything they normally do, except that everyone’s a p-zombie. That’s easy for an epiphenomenalist to imagine, because our own world is basically that plus some extra stuff beyond physics, caused by the physical world but never causing anything in it.

I find this so strange that I can’t entertain it. Smarter people than me do entertain it, so take my reaction with a grain of salt. But after 15 years of wrestling with the idea it remains the theory of phenomenal consciousness I give the lowest odds of being true.

As a well-behaved physicist, I’m also very hesitant to cut the first option. There are lots of other problems with cutting it. Evolution gives us a lot of reason to think we’d see ourselves and other people as monumentally important compared to the rest of the world, so it’s a little suspicious that the one place our most successful scientific theory breaks down is inside our own heads.

I can see myself believing this if I were religious. If you are religious, and your belief that the mind transcends physical reality is why you don’t believe machines can think, my only ask is that you say that directly in conversations about it. I often find that debates about minds and machines can meander in meaningless details for hours before one person flags “Oh by the way, the main reason I’ve been denying your position is my religious beliefs, which I’m only mentioning now.” That clarifies where the actual disagreement is.

If you’re not religious, there are a ton of additional problems with cutting Option 1. Evolution by natural selection is a purely physical process. How did merely physical animals at some point start accessing, affecting, and being affected by some ethereal realm or substance beyond physics? Why are they the only entities in the universe that can access this? If you cut the first option, you’re not religious, AND that’s why you don’t believe AI can think, I think the burden of proof is squarely on you to explain how your view fits with a broader scientific picture of the world.

If you do cut the first option, an interesting question is how much of our thinking and understanding happens outside the physical world. Do we have thoughts that happen outside of our brains? If so, why does severe brain damage also impair specific kinds of thinking? Whatever role the nonphysical part of consciousness plays, our slow, deliberate, higher-level thinking still depends massively on the physical brain working in specific ways. That leaves open the possibility that machines without the inner spark of phenomenal consciousness could think this way too.

The last option to cut is 2, where we say that the sense that consciousness is somehow above physics is some kind of complex illusion. This is the one I lean toward, but only because the other two look so much worse. There’s a lot that troubles me about this view, most of all:

Come ooooon! It really really really really seems like there’s something happening here that isn’t just information processing. There’s a “way it is like to be me” that can’t be articulated with sufficient physics equations. Aaaaah!

But phenomenal consciousness is incredibly strange. There are ways it doesn’t cohere at all with our folk theories of it. It’s strange enough that I can at least entertain the idea that it’s a complex misunderstanding happening in an information-processing machine called my brain. To build your intuitions for how this could be possible, I’d very strongly recommend Dennett’s paper Quining Qualia, one of the most fun and accessible papers in philosophy of mind. Another good essay is Joe Carlsmith’s Grokking Illusionism.

This view, that phenomenal consciousness is reducible to physical processes, has many flavors, like illusionism. However weird you find this idea, epiphenomenalism is just as weird, and it adds a bizarre extra entity that doesn’t behave anything like how I experience consciousness. For example, whatever we mean by consciousness, I personally mean the thing that causes me to see that red box and type the words “Aaaaah!” But epiphenomenalism tells us my typing that was not at all caused by the redness of the red.

There are also philosophers who, when faced with this trilemma, throw up their hands and announce that our concepts of the world are simply not built to understand consciousness, and the answer here might be stranger than we can conceive.

How does any of this matter for whether AI could at some point think? Well, which story we believe tells us something about how phenomenal consciousness interacts with our thinking.

If you believe in epiphenomenalism (that phenomenal consciousness exists but has no causal power), then our having phenomenal consciousness has never affected the world at all. We could lack it and the world would look exactly the same. And humans can think and understand things. So on this view, whether AI could think and understand doesn’t depend at all on whether it has phenomenal consciousness, because our having it doesn’t matter to our thinking either.

If you believe consciousness is nonphysical and does affect the physical world, I’ve said why I think the burden of proof is on you, and even then most of our deliberate thinking still seems to run on the physical brain.

If you believe (like me) that phenomenal consciousness is probably entirely reducible to physical interactions, then whatever role it plays in thought could in principle be replicated by other physical processes. This is necessary but not sufficient for showing that AI could have it. Consciousness could be a special kind of physical process that depends on the substrate (the material the mind runs on). This view is currently defended most prominently by the neuroscientist Anil Seth. But Seth (like many others who hold this view) believes intelligence and consciousness can come apart, so machines could “think” and “understand” without being conscious. The question for people who hold this view becomes whether machines can do the correct type of physical things that human brains do, not whether there’s some extra-physical spark inside the machine.

If I consciously notice a mistake, and that lets me correct an argument or line of thought, something has changed in how my brain is processing the argument. Relevant information has been successfully introduced. We can look into what’s changed, and how it’s improved my thinking, seemingly entirely via information processing, without ever explaining what it felt like to notice it. If phenomenal consciousness is not reducible to access consciousness, this makes it look like one specific way information is introduced to my mental system, and anything it introduces that can’t be fully reduced to information (the redness of the red etc.) doesn’t seem to play a direct role in thought.

The fact that “I am experiencing the redness of red,” can be reducible to information the brain can use (I can write it out as a sentence that you can read and make sense of) but the phenomenal content remains out of reach to summary by pure information. But this same phenomenal experience also becomes instantaneously unavailable to our own minds the second our perception changes. If I look away from the red object, it seems like the only way I can use that past phenomenal experience is via the information I’ve drawn from it that I can record. “I saw a red object, it has the same wavelength of light as other red objects. Also, I saw a deep redness of red. Aaaaah!” It seems like thought and understanding cannot depend on what we’re phenomenally experiencing at any one time. I can form thoughts about the color red without seeing it in the moment. And when we’re not experiencing something directly, it seems like the only access we have to it is via the third-person information it provided (our past selves are effectively other people here). So unless there are types of thinking that we can only as we’re having the phenomenal experience in the moment, it seems to me like the only parts of consciousness that can affect thinking are reducible to access consciousness. That means that for thinking about whether AI models can think, we need to turn to theories of access consciousness. Whether AI has phenomenal consciousness thus looks to me to be pretty unrelated to whether it can think.

Access consciousness’s role in thinking

Consciousness is a controversial topic. A recent attempt to map all theories of consciousness lists over 200. As someone who’s far from an expert myself, I can’t confidently argue for any one. But only a few theories get most of the serious academic attention, and they all contribute to my case that AI can think. I’m going to summarize what each says about the role of consciousness in thought, and how it compares to the folk “movie theater” picture of the mind. In every one, processes we’re not conscious of do the actual work of thinking. Consciousness is at best a way of coordinating those processes, or something separate from them that’s helplessly along for the ride, just justifying their results. A few of the most prominent theories imply that everything we’d call thinking could happen with no consciousness at all.

Global workspace theory

Global workspace theory is probably the most popular scientific theory of consciousness right now. It describes the brain as full of specialized processes all running at the same time, almost all of them unconsciously. Some processes do things like recognize faces, or enforce rules of grammar, or keep you physically balanced. Each one is good at its own narrow job in the broader system, and they mostly cannot communicate with each other directly. But they share a general “workspace” that can only hold a little at a time, so different pieces of information compete to get into it. Whatever wins gets broadcast to all or most of the processes in the brain at once, so the whole brain can use it. According to this theory, being conscious of something is just the information being broadcast.

The theory’s creator, Bernard Baars, ironically used a theater metaphor to describe it. The workspace is the stage, attention is a spotlight on the stage showing what’s on it, and the audience is the giant crowd of unconscious processes in your brain. In this picture there’s no self, and no single you in the audience watching the whole thing. Baars himself said “You don’t have a little self sitting in the theatre.”

Some neuroscientists (most prominently Stanislas Dehaene) have tried to translate this into a theory about physical brain tissue, where the workspace is a network of neurons with long-distance connections. When you become conscious of something, activity suddenly spreads across that network.

If the theory is right, individual mental processes should each have information the others don’t, and we only become conscious of it when it’s broadcast to lots of them at once.

The first part of this claim is pretty self-evident. One of my favorite tests for this is to think of a word at least 5 letters long, and then (without looking!) try to picture where each letter is on a standard QWERTY keyboard. I personally find this is very hard. Next, place your hands on a hard surface like a table, and mime typing out the word as if you had a keyboard below you. When I do this, my fingers show me exactly where each letter is. The muscle-memory part of my brain had this information all along, and wasn’t sharing it with the part that forms visual memories.

Scientists have come up with ways to try to test this theory. In one 2009 study, researchers flashed numbers on a screen too quickly for people to consciously see them. People could still do slightly better than chance at simple operations on the mystery number, like naming it, saying whether it was bigger than 5, or subtracting 2 from it. But they failed at chaining these operations together (like subtracting 2 and then saying whether the result was bigger than 5). This suggests individual unconscious processes picked up bits of information about the number, but couldn’t combine what they had, because the number never got broadcast to the whole brain. A similar pattern will appear in AI models in Part 2.

This theory stands out for giving consciousness a clear, active role. But even here consciousness is more of a spotlight than a director of what’s happening on the stage. It only needs access consciousness, with no transcendent viewer guiding things from above. This also doesn’t rely on the specific nature of neurons, because a machine could just as easily broadcast information to lots of individual processes.

Higher-order theories

Higher-order theories of consciousness say that a mental state is conscious if your brain is also representing itself as being in that state. A red apple might be in your field of vision without you being conscious of it. You only become conscious of it when another part of your brain forms the belief that you’re seeing the red apple. The philosopher David Rosenthal is the best-known defender of higher-order theories.

Blindsight, a bizarre condition, seems to fit this theory well. Some patients with damage to their primary visual cortex report that they can’t see in part of their visual field, but do significantly better than chance when asked to point to things in front of them or guess basic shapes. A higher-order theorist might explain this as their brains still processing some visual information without the higher-order representation required for conscious seeing.

These higher-order stories the brain tells can easily be inaccurate. Worse, they don’t seem to play much role at all in most of what we normally consider our conscious thought and decision-making, like reasoning or planning. Rosenthal argues that reasoning and planning are done by other mental processes whether or not they’re represented in consciousness, and their becoming conscious “adds no significant function.” So under this view, thought and understanding do not seem to depend on consciousness at all.

Attention schema theory

Attention schema theory (AST) describes the brain making lots of important but unconscious decisions about where to spend its limited processing power. Giving a process more of that power is what it means to give it “attention.” If you’re trying to follow one person in a crowded room full of other conversations, your brain can give much more processing power to their words. Michael Graziano, who came up with AST, points out that your brain keeps a simplified model of your body to control your movements (reaching for a cup means making a lot of assumptions about where your hand is). He proposes that it also keeps a simplified model of its own attention, to control where that attention goes. Among other things, it would be bad to let anything that shows up in your senses grab a lot of processing. Your conscious awareness is this simplified model your brain runs to manage itself. Here consciousness plays a useful role, giving other brain processes information they can use to manage how much processing goes to each task.

Introspection under attention schema theory is also a simplified model of what’s happening in our brains. Again, our consciousness doesn’t have direct clear access to what’s happening in our own heads. Your sense of yourself as an indivisible observing entity is actually the brain telling a useful simplified story about itself.

Graziano thinks that AST implies that a machine with a similar attention model of itself could form the same conviction that it has subjective awareness. So this is another theory where access consciousness plays a useful role in the brain that could be replicated in machines.

Predictive processing

Predicting our incredibly complex environment was central to our ancestors’ survival, as it is for other complex animals. Predictive processing approaches make this central to explaining conscious perception and how the brain works. Here, your brain is constantly predicting sensory input using internal models shaped by its experience. Prediction concerns both what is causing your present sensory input and what future input may arrive. The brain compares these predictions to incoming sensory signals and updates its internal models to better match them. On this account, your conscious perception is your brain’s best interpretation of sensory signals. It’s the brain’s best predictive guess about the world. In this theory, you don’t draw raw information into your conscious mind and then interpret it. Your conscious experience is the interpretation of the raw data itself. A lot of our everyday experiences can be explained this way. The reason you usually can’t tickle yourself may be that your brain predicts the sensory consequences of your own movements and dampens its response. Looking into a hollow mask from the back can create the illusion that it’s a face bulging out at you instead, partly because our brains strongly expect faces to protrude outward rather than cave inward.

This is remarkably similar at a broad level to the basic idea behind modern language models, where by “merely predicting the next word” models develop complex and subtle internal representations during training that enable them to write coherently about the world. I’ll flag here though that Anil Seth is a prominent supporter of predictive processing as an explanation for what happens in the brain, and he’s also skeptical that current approaches to AI will produce consciousness. Among other things, he suspects that consciousness is too tied up with the actual biological processes that keep us alive for computation alone to reproduce it. He accepts that machines can be very intelligent and separates consciousness from intelligence. He also entertains the possibility that machines could understand without being conscious, though he remains cautious about whether current language models do. My own take is that this theory sounds strikingly similar at a broad level to descriptions of large language models, to the point that while I accept that they might not be conscious, taking predictive processing seriously as an account of perception would make me more open to thinking that current AI models can think and understand.

Integrated information theory

I’ve always found integrated information theory (IIT) strange, but enough smart people are into it that it deserves a mention, though I’ll note that in 2023 over 100 researchers signed a letter calling it pseudoscience.

IIT’s starting observations include that experience is unified (you experience the parts of your visual field together, as one experience), and specific (you have one exact experience out of countless possibilities). What physical system could have both of these properties? Maybe something with parts that influence each other in such a way that the system as a whole is irreducible. Dividing it into independent parts changes the cause-and-effect structure of the whole thing.

As an example, the system underlying the human visual field can’t, on this view, be like an idealized camera’s sensor array. The array is made up of individual sensors that don’t notice or react when others are on or off. You could divide that array in half without changing how each half works. The brain instead contains dense loops of mutual influence. IIT tries to quantify integrated information in a system’s cause-and-effect structure with a score (called “phi”), which, according to the theory, measures how conscious a system is. Consciousness, on this view, is specific, irreducible cause-and-effect structure within a system.

One interesting fact that lends some credence to this view is the fact that the cerebellum has about 4 times as many neurons as the cerebral cortex, but damages to it seem to not affect consciousness nearly as much as the cerebral cortex. The cerebellum has a lot of relatively independent processing units, which according to IIT give it low phi.

This has a lot of weird implications, that make me think it’s probably not correct. The computer scientist Scott Aaronson pointed out that IIT’s method of calculating phi could make simple logic-gate arrangements that don’t do anything interesting or useful in the world more conscious than humans. The creator of IIT does accept that a sufficiently large grid of logic gates could be conscious even while it’s inactive. That seems pretty bizarre, but consciousness is bizarre, so here we are.

My read is that IIT is specifically about phenomenal consciousness. I’m not sure what role access consciousness plays in the theory, and I can infer that it doesn’t play much at all. As an example, a one-way network without any loops or dependencies has zero phi. And any network with loops can, in principle, be rebuilt as a one-way network that produces identical outputs. This seems to mean that the theory implies that anything and everything done by a conscious entity could also be done by an entity with zero consciousness. A recent paper using IIT as its framework seems to agree that consciousness and all abilities of the brain come apart, and so AI can do everything the human brain does without being conscious. Thus, IIT does not rule out the possibility of thought and understanding in AI systems.

These all seem to imply that the aspects of access conscious important for thought could exist in machines

In all of these theories, most of the work of thinking happens outside consciousness, and where consciousness does help, it’s by sharing information or managing where processing power goes. If that’s consciousness’s whole role in thought, it seems like something we could easily imitate in AI. None of this requires the AI to have phenomenal consciousness, only (sometimes) access consciousness. And if access consciousness is mainly about managing the flow of information (sharing it across a system, and deciding how much processing power different parts of the mind get), that sounds like something a machine can do. So even if you’re very skeptical that machines could ever be conscious, that tells you very little about whether they can think.

Any part of your mind that understands something is made up entirely of parts that don’t understand

This one seems too obvious to need saying, but it helps keep me grounded in the debate, and a lot of people talk as if it isn’t true.

Unless thought and understanding are supernatural and happen outside of physics, any system that understands or thinks is ultimately made up of fundamental particles and forces. These on their own don’t understand or think. Thinking is an emergent property of arranging them in specific ways to perform specific functions. At a higher scale, individual neurons also don’t “think” or “understand.” They are building blocks for a specific system that thinks and understands. All that human neurons do is receive basic chemical signals, change their electrical activity, and release some chemicals. Any system that understands will be made up of parts that “only” do simple things. That’s a basic part of living in physical reality.

It’s surprising how often people will say other systems can’t think or understand because they are made up of individual parts that don’t understand. People will say that an AI is “just doing matrix multiplication,” as if the fact that none of those individual mathematical operations understands anything means the whole system cannot understand anything either.

The role of access consciousness in thought, like anything else in the world, also needs to be completely explained by individual processes that are not conscious. If thinking requires some higher-level awareness, that awareness also has to be explainable in the language of physics (again, unless you believe it happens in a separate nonphysical world that interacts with ours). Dennett once observed (in his book Brainstorms) that explaining an intelligent system means breaking it into smaller, dumber systems, then breaking those into even dumber ones, until you get down to parts so simple that they can be “replaced by a machine.” I think this is the only way minds can ever be explained.

Language, meaning, and consciousness

The folk picture also leads people astray about what it means to understand language. The Cartesian theater picture of language is that to understand a word, you need a clear, first-person experience of what the word refers to. Anything without this experience doesn’t really understand the word. According to this theory, when I say “apple,” you understand me because you can summon a first-person image of an apple, or at least have experienced one before. AI might have a huge collection of associations with “apple,” but without an inner Cartesian theater to experience what the word refers to, it can never truly understand it. So (the theory goes) AI models only manipulate words, instead of using and understanding them the way we do.

Again, I’ll start with some first-person experiments you can run on yourself. Think about the word “dog.” Maybe a picture of a dog appears in your mind. This passes the folk test of understanding. Now think about the word “unless.” What appears in your mind when you do that? You obviously understand “unless” perfectly well. But it’s hard to see what clear inner experience could contain the meaning of “unless.” “Unless” changes the relationships between other things you could imagine, and mental pictures don’t seem much use for a word like that.

So maybe we can understand relational words without mental pictures, but simple nouns do require mental pictures. Already, this seems like a weird arbitrary distinction. If you can fully understand a word like “unless” without a mental picture attached, this suggests mental pictures aren’t that fundamental to understanding. Another problem with this view is that some people have aphantasia and can’t form visual images at all. These people can go most of their lives without even noticing that their mental experience is weird or different in any way. If understanding words required mental images, these people would regularly use language incorrectly, or not really understand the words they use. That seems obviously wrong.

You might object that these people have still seen the things they talk about. They’ve still seen dogs even if they can’t form pictures of them. But there are obvious counterexamples. Do blind people not understand the language they use? Do I not understand the word “Beijing” because I’ve never been there? At this point, I worry that when people say “understand” what they mean is “know everything there is to know about” something. I don’t know everything there is to know about biology, and I can’t form a clear, complete picture of it in my mind, but I can use the word, and I feel like I deeply understand what it means. This understanding is, I find, partly inaccessible to my consciousness. There are a lot of details about what “biology” means that I can’t hold in my conscious mind all at once.

So it seems like phenomenal conscious experience can give us new information about something. It can show me what it’s like to experience the red of a sunset. But a blind person who can’t experience this obviously still “understands” the word “red.” They can use it coherently and know how red relates to lots of different things in the world. They’re just missing one specific aspect of it: the phenomenal first-person experience of it. This seems more like me using the word “Beijing” without having been there than me using the word “Buchladen,” which I don’t know the meaning of at all. Phenomenal consciousness does not seem to be required to completely understand the meaning of the words we use.

This gets at a confusing way we use “understanding” to mean two different things. It can sometimes mean a deep first-person knowledge of what an experience is like. If someone told me that I don’t understand what it’s like to live in Beijing, I’d agree. There’s a lot about the day-to-day experience of living there that I don’t know. But if someone said that I don’t understand the sentence “what it’s like to live in Beijing” that’s obviously wrong. I know what the words are referring to and how to use them. In conversations about AI, these two definitions are sometimes blurred. When I say “AI understands the words I’m saying” I worry that what people think I mean is that AI has some first-person subjective experience where it deeply relates to me. All I actually mean is that AI has enough information about the words, and their relationship to all the other language and images it can process, that it can work with the information in them about as well as I can.

My own view is that understanding a word is a matter of what it lets you do with the word. Consider what it means to understand “fragile.” Someone who really understands “fragile” knows that dropping a fragile object is risky, and that whether something is fragile often depends on what it’s made of. They can use the concept in new situations. To me this looks like having a lot of specific, hard-to-articulate information clustered around a concept like “fragile” (including the logical relationships it implies, like that fragile things are hard to ship), which often can’t be boiled down to a single clear idea we can hold in our heads. Our conscious experience is just one way of getting closer to the right cluster of information and relationships, and it isn’t what holds that cluster together.

One role language plays in thought is letting us understand things we’ve never encountered by connecting them to things we already understand. I can explain a tool you’ve never seen by describing what it does and what it’s made of (all things you already understand), and you can understand the new tool without a clear mental picture of it or any first-hand experience of it.

A problem in the philosophy of language often brought up in discussions about AI is the “grounding problem.” Our language can’t just be a web of relationships between words. It has to be grounded in the reality it describes, so the relationships between words need to be determined by the relationships that actually exist in the world.

This is an interesting problem, but the folk theory often leads people to believe that any solution to the grounding problem requires a first-person conscious observer seeing for themselves “what the language is about.” They need to have the phenomenal experience of the world and connect that to the words they use. If we drop the folk theory, our observations become basically a way of collecting a lot of information about the world at once. As long as information from the world shapes what language says about it, well enough that predictions made in language match what’s actually happening, that seems completely adequate for solving the grounding problem. We don’t need floating observers from outside physical reality giving words meaning based on subjective experiences that information processing can never reach.

The grounding problem does mean that information processing alone isn’t enough for language. A being could have a perfectly coherent internal language that does not at all relate to anything in the world. The main question for whether something “understands” its words is whether they’ve been shaped enough by information from the outside world to reliably describe and predict it. I’ll argue in Part 2 that AI has several ways of getting this information, mainly during training.

So once we drop the folk theory, understanding language means being able to use words correctly by other speakers’ standards, make a huge number of subtle, correct connections between each word and other words, use words shaped by the outside world to correctly describe and predict what happens in it, apply them in unfamiliar situations, and correct mistakes. I think consciousness is basically one way humans acquire the information that structures our language, not the thing that gives our words meaning. If something has all of these properties, I would say it understands words for the same reason that humans understand them.

Conclusion

Here’s a summary of the picture I’ve tried to paint here:

  • A lot of the work of thinking happens before we become aware of the results. Whatever story we tell about thought and understanding needs to include the fact that a lot of it is capable of happening outside of our conscious attention.
  • Our inner world is often opaque to us. It is easy to quickly develop incorrect beliefs about our very recent conscious experience and decision-making.
  • Introspection is just another fallible mental activity that draws on information from the world and past mental activity, like anything else. We don’t have some inner god’s eye view of our mind where we can clearly see everything about our own thought process. Perfect self-knowledge can’t be a requirement for AI thinking, because humans can think without it.
  • Careful reasoning depends on specific complex kinds of coordination, like holding onto intermediate results, brining the relevant information together, understanding the semantic content of our words, notice conflicts, direct attention to different places, and revise mistakes. These are substantial, but none of them currently seem inaccessible to AI models. Without a Cartesian theater guiding them from above, this look much more accessible to machines.
  • It seems like whether AI has phenomenal consciousness does not matter for whether it can think. What matters is whether it has the specific qualities of access consciousness that facilitate thinking. The best theories of access consciousness all seem to imply this is possible in principle. Having an experience and being able to work with all the information that experience would provide seem like different things, and thought and understanding seem to exclusively depend on the second.
  • We don’t actually have to settle every debate and disagreement about consciousness before taking machine thinking and understanding seriously, because many of the most popular accounts of both phenomenal and access consciousness imply that AI could think and understand with the right kind of information processing.
  • The parts of a system do not need to understand anything for the whole system to understand. Individual neurons don’t understand what’s happening. The argument that a system cannot understand because it’s made of fundamental building blocks that don’t understand is confused. Everything physical is made of these building blocks, so any story of systems that understand has to build them ultimately entirely of things that don’t understand and could be “replaced by machines.”
  • Understanding words requires knowing what follows from using them, which is structured by how they relate to the outside world, and their broad logical and semantic connections to all other language. It does not require clear coherent phenomenal conscious mental images.

Once we poke at the folk conception of our minds, the picture that comes out is a mind with abilities spread across many interacting process. Some receive information, or find patterns, or catch mistakes, or generate possibilities to investigate, or follow logical steps from premise to conclusion, or broadcast information to lots of other parts. Our stories we tell about our own minds are products of these imperfect piles of capabilities working together in complex ways, aided by what we’d call access consciousness, which itself appears to be a pile of specific information processing abilities. It’s amazing what these all can do together.

This all I think radically changes what we should expect from machine thought. When investigating whether a machine can think, this gives us way more specific questions to ask, like whether it can represent the world with models, whether it has information about the subtle semantic content of the words it uses, whether it can monitor the output of its past processes and look for any problems or contradictions, whether it can follow a trail of logic from premises to conclusions. Asking “but is it human?” after all this makes less sense to me. It seems like asking whether a submarine is a fish after establishing whether it can dive, navigate, and cross an ocean. These are all individual qualities that make up but don’t completely summarize both fish and submarines, and confusing them as being inseparably tied together in either is obviously silly. We have strong evolutionary reasons to treat the human mind as special, but once we poke at what we actually know these abilities also look like amazing but separable abilities and functions that could with the right technology be replicated in machines.

I haven’t yet given you any reason to think that current AI can do any of this, that’ll be in Part 2. I hope I’ve shown why “the AI doesn’t have conscious experience” can’t end the discussion about AI thought. By itself, the claim gives us very little reason to think that a system cannot do most or even all of what we mean by thinking. I believe that the main question facing us now with AI is whether AI will be able to do all the types of thinking human minds can do, or even valuable forms of thinking we’ve never accessed, and what the world looks like when they can. The difference looks way more like one of degree rather than kind to me.

I worry that people who want to enforce strict bans on “anthropomorphic” language being used to describe AI are often implicitly enforcing the folk conception of how our minds work. They do not want to pick apart the mind’s cognitive abilities and list which AI has and does not have, because for them the human mind seems so obviously special and transcendent that it’s somewhat sacrilegious to imply it could even be divided up into these individual abilities, cutting them off from the unifying light of phenomenal consciousness. This enforcement of a ban on anthropomorphic language is often done in the name of cold hard science up against superstitious suggestible people who have mistakenly identified a ghost in the machine. In reality, I’m worried the people enforcing it have superstitiously discovered a ghost in themselves that isn’t there, and have closed themselves off to a serious attempt to move beyond our own self-concepts and demystify the surprisingly inhuman processes and forces that underly our own minds to the point that some of those processes could be considered to have already been achieved in machines.

How this could all be wrong

I’ve painted a picture here that’s very sympathetic to physicalism, the idea that everything (including consciousness) is physical. Physicalism could be wrong, and that’s the main way I think my story could fail. In polls of professional philosophers of mind, physicalism is more popular but far from a consensus in the field.

It might be that phenomenal consciousness has nonphysical aspects that can’t be reduced to access consciousness, that these aspects somehow depend on the material the mind runs on (brain matter vs. computer chips), and that thinking and understanding depend on them way more than I’m giving them credit for. I think the odds of this are low, but not too low, so I’m open to the idea that machine thinking will be permanently limited by lacking this nonphysical ingredient. But my default assumption is still that this is all wrong, and that machines can think.

The other way I could be wrong is if the material or structure of brains lets them do things with information that computers, given their design, never can. That would mean computers are more limited in how well they can mimic access consciousness. This is interesting, but it doesn’t seem likely to me that computers will be so limited they can’t mimic most of the important stuff in thinking. I think the case that they aren’t gets much stronger once we look at what current AI can actually do, and why, which is where Part 2 will start.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论