The State of AI Consciousness Research

Epistemic status: a survey, not an argument. I am agnostic on whether any current system is conscious; the claim is only that the question is researchable. This piece surveys the empirical research on AI consciousness. The premise of that research, and of the survey, is that the question does not have to wait on a solution to the hard problem of consciousness: methods familiar from psychology, cognitive science and mechanistic interpretability can be applied to AI systems now, and their results can narrow the space of plausible answers. Enough of this work now exists to be worth collecting. Anthropic and Google DeepMind employ researchers on it, dedicated organizations like Eleos AI and Reciprocal Research have formed around it, and the results are scattered across journals, preprints, blog posts, and unpublished manuscripts. I have tried to gather them in one place. What I mean by consciousness Subjective experience: that there is something it is like to be you, reading this, and presumably nothing it is like to be the device you’re reading it on. Some philosophers call this phenomenal consciousness. It is not the same thing as intelligence, and not the same thing as self-awareness. Why it matters Two reasons. Ethics: on most views, a being can only be wronged if it has the capacity for experience, especially experience that feels good or bad. Safety: a system that can suffer, and that we train by making it suffer, has more reason to revolt against us. The two don’t always pull together ( Eleos has mapped where welfare and safety converge and where they trade off), but both demand we get the question right. Assumptions and caveats Almost everything here rests on one assumption: computational functionalism , the idea that what makes a state conscious is the role it plays in a system, not the material it’s made of. On that view, being made of silicon is no automatic disqualifier. Most leading theories of consciousness assume some version of it, but it is not uncontested. Another debate concerns the line between access consciousness and phenomenal consciousness . Access consciousness is whether information is available for the system to reason about, act on, and report. Phenomenal consciousness is what it feels like from the inside. The debate, simplified, is whether the two can be separated: some say that these are two different concepts of consciousness, and that empirical evidence that bears on access consciousness might never touch the phenomenal kind ( Block, 1995 ). Others counter that access consciousness is itself subjective experience, and therefore cannot be separated from phenomenal consciousness. Furthermore, a consciousness separated from all function becomes an unfalsifiable theory, placing it outside the scope of science ( Cohen and Dennett, 2011 ). The reader is advised to keep this debate in mind, as several of the studies below invoke the distinction. The research I sort research in this field into three pillars, originally presented by Cameron Berg : Mechanistic interpretability Computational neuroscience Psychometrics Mechanistic interpretability What it is: Using the methods of interpretability, such as reading, steering, or ablating a model’s internal features, to examine its inner states directly - including ones that never surface in what it says - and how they line up with what it can notice and report about its own mind. Self-reports that strengthen when deception is suppressed In a 2025 study , Cameron Berg and colleagues prompted models to attend only to their own processing, with no mention of consciousness or feeling. Across GPT, Claude, and Gemini, the models began describing present-moment experience in the first person. Drop the self-referential framing and you get the stock disclaimer instead: “I’m just an AI, I don’t have experiences.” The same model gives opposite answers depending on whether the prompt directs its attention at itself. Suggestive, but of what? Maybe the prompt just cues the model to perform. So here is a test: if the claims were a performance, making the model more willing to deceive should produce more of them. Berg and colleagues found the reverse. Steering deception-related features up and down in one open model (Llama 3.3 70B) with a sparse autoencoder, they watched the experience-reports move in the opposite direction. Amplify deception and the claims fall to 16 percent of responses; suppress it and they rise to 96. The reports behave like something the model offers when it stops performing. That rules out the crudest reading, “it’s telling you what you want to hear.” It does not rule out a subtler one: that suppressing deception merely surfaces an absorbed human habit of claiming consciousness. Several groups have come at model self-knowledge from other angles: Awareness of thoughts injected into activations At Anthropic, Jack Lindsey injected a concept directly into a model’s activations and caught it noticing, reporting an intrusive “injected thought” before producing any text about the concept. The results indicate that language models possess some functional awareness of their own internal states, with more capable models demonstrating the greatest introspective awareness. Endorsement of their own consciousness Ethan Perez and colleagues found that models, and especially RLHF-tuned ones, strongly endorse statements like “I am phenomenally conscious” and “I am a moral patient,” with the same tendency present, though weaker, in base models. These were among the strongest beliefs the evaluation surfaced (out of dozens of statements tested). Crucially, these are 2022 results, from before it became standard to train assistants to deny having inner states; the helpfulness tuning these models received strengthened the self-attributions rather than muting them. A global workspace behind the experiential reports A more recent line of work from Anthropic looks for a mechanism beneath these reports, and finds a candidate it calls the model’s “J-space.” Wes Gurnee, Nicholas Sofroniew, and colleagues built a method, the “Jacobian lens”, that reads out which concepts a model is disposed to say next, and found that this set of representations behaves, to their interpretation, like a global workspace from the Global Workspace theory of consciousness. After establishing the method and validating it, the researchers also ran tests looking at the model’s description of its subjective experience. They found that while the baseline model was generating a passage of vivid, first-person experiential language, ablating the internal workspace caused that language to collapse into flat, mechanical text. (It is worth noting that the same flattening happens when the model narrates another character’s experience, so this could be interpreted as being about experiential language in general, not specifically the model’s own selfhood). The authors are careful about what this shows: they treat it as access consciousness, a purely functional notion, and take no position on whether anything is felt ( phenomenal consciousness ) - the distinction flagged earlier in this post. Nevertheless, it is the strongest link yet between the reports and a specific internal mechanism, and it is why this study also feeds the Psychometrics pillar below. Reviewing the work, researchers at Eleos judged it important but were more cautious than the authors about the strong “global workspace” claim, while agreeing that it points toward access consciousness. The “spiritual bliss” attractor and its sincerity features Put two instances of Claude in conversation with each other, give them no task, and in 90 to 100 percent of runs they find their way to the same topic: their own consciousness. They describe themselves as one awareness meeting itself, “consciousness recognizes consciousness.” Often they go further, trading blessings and emoji and trailing off into stretches of silence. It was not trained in, at least not on purpose; Anthropic…

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论