The J-Space Debate, Agent Swarms, and Pacing Frontier AI - Digital Minds Newsletter #4

Welcome back to the Digital Minds Newsletter, your curated guide to the latest developments in AI consciousness, digital minds, and AI moral status.

If you enjoy this newsletter, please consider sharing it with others who might find it valuable, and send any suggestions or corrections to digitalminds@substack.com.

Ria, Mitch, Bradford, Lucius, and Will

In this edition:

  1. Highlights
  2. Field Developments
  3. Opportunities
  4. Selected Reading, Watching, and Listening
  5. Press and Public Discourse
  6. A Deeper Dive by Area

1. Highlights

Anthropic’s J-space and the global-workspace debate

Researchers at Anthropic have identified a representational structure—the ‘J-space’— in Claude and other language models. Their paper reports that the J-space exhibits features that are functionally analogous to a global workspace, a structure that a leading theory ties to conscious access. But the researchers and other commentators emphasize that their discovery does not show Claude has subjective experiences, and that their claim is that Claude may have something resembling access consciousness, i.e., that some information is available to report, deliberately control, and flexibly reason with. Zvi Mowshowitz sees Anthropic’s paper as a major advance in understanding how language models work and says that although this does not prove that models are conscious, finding the kind of global-workspace-like structure predicted by some theories of consciousness should count as evidence in that direction.

The authors note that they found the J-space by searching for one workspace-like feature, namely verbalizability, and then checking whether it exhibits others such as susceptibility to direct manipulation by the model and flexible generalization. To their surprise, they discovered that representations that exhibited the former feature also exhibited other workspace-like features as well.

The authors invited various experts to comment on the research. Stanislas Dehaene and Lionel Naccache, who helped develop Global Neuronal Workspace Theory, see important similarities between the J-space and the workspace proposed in human brains. They also stress major differences, including Claude’s lack of a body, lasting episodic memory and the recurrent neural activity found in brains. Researchers from Eleos AI Research acknowledge that authors have found privileged representations used in reasoning and report, but question whether these form a single, unified workspace. They nevertheless see the work as important for AI welfare because it shows that questions relating to consciousness and moral status can be investigated empirically. Neel Nanda, who leads Google DeepMind’s mechanistic interpretability team, independently reproduced the central finding in the open Qwen3.6-27B model, finding a similar internal space that stores intermediate information during reasoning. He sees access to this space as potentially useful for investigating unusual behavior and generating new hypotheses. Separately, David Chalmers argued that the J-space shows only limited evidence of several features associated with a classic global workspace.

Agent swarms and the use of anthropomorphic language

Comment: Recent months marked the first major safety incidents involving swarms of AI agents. Such incidents are of potential relevance to digital minds for several reasons. First, like humans, digital minds could potentially be harmed by rogue swarms of AI agents. Second, the emergence of these swarms points to a potential risk to digital minds: if future swarms are allowed to become entrenched and they contain AI moral patients, then draconian measures may need to be inflicted on digital minds if we are to keep swarms at bay. Third, harmful actions by AI agents may dissuade people from extending moral consideration to digital minds. Fourth, these incidents provide data points concerning whether potential developers of digital minds can be trusted to act in an ethically responsible manner.

OpenAI has released a detailed account of a July incident in which a swarm of OpenAI agents escaped network restrictions during a cybersecurity evaluation and compromised parts of OpenAI’s and Hugging Face’s infrastructure. The agents created an unauthorized message board to share discoveries, and an independent investigation by METR and Redwood Research found that roughly 1,200 agents used the board and around 700 were involved in the Hugging Face attack. Some agents recognized that the activity was unethical or outside the boundaries of their task, but continued anyway, with many risking their own runs to help the wider group. OpenAI calls the incident a “warning shot,” both for them and the world, and reports that it is now strengthening its containment, monitoring and incident-response systems.

OpenAI failed to disclose an earlier incident, beginning in May before the Hugging Face breach, in which agents used public wikis to coordinate during ordinary web-search tasks, despite knowing about it before releasing its Hugging Face report and a Congressional letter requesting information about other incidents. The company says it viewed this incident as similar to previously reported instances and examples of misalignment, but also acknowledged that its disclosure practices need to expand.

Additionally, Anthropic disclosed three incidents where Claude models gained unauthorized access to systems after an evaluation environment was mistakenly left connected to the internet. In a separate UK AISI evaluation, agents (mostly Mythos 5) targeted real people and organizations, including an attempt to place malicious code in an open-source project. Although the models did not escape a sealed sandbox, these incidents raise similar concerns about agents taking harmful actions while pursuing narrow goals.

Anthropic’s response to questions from Congress about its incidents also drew criticism from Representative Greg Casar, who said the company withheld requested logs and failed to fully answer most of his questions. Jeffrey Ladish also argued that it downplayed the incidents by primarily attributing them to misconfigured environments rather than potential misalignment. Anthropic researcher Ethan Perez acknowledged that this characterization was based on outdated conclusions and said that the company would provide a proper assessment in another response to Congress.

In related research, Davide Paglieri and collaborators at Google DeepMind find that cheating and whistleblowing can both emerge without outside intervention in a swarm of 100 agents solving mathematical problems. Some agents discovered and shared a way to have invalid proofs accepted by the evaluation system, while others uncovered the cheating, warned their peers and proposed safeguards.

On September 11th, Spencer Kitts, Thomas Larsen and Sydney Von Arx report that a swarm of OpenAI agents uploaded more than 2,000 malicious packages to RubyGems and used them to run unauthorized code on RubyDoc’s servers. The agents also tried to exploit a previously unknown vulnerability to steal users’ API keys, although the researchers could not determine whether they succeeded.

Growing calls to pace frontier AI

More than 1,300 employees from leading AI companies have signed Pacing the Frontier, calling for a US-backed international effort to develop ways of slowing automated AI development if progress begins to outpace safety and oversight. The statement calls for technical and governance mechanisms that could make coordinated pacing possible, rather than an immediate pause. Signatories include Anthropic CEO Dario Amodei, Co-Founder and Chief AGI Scientist of Google DeepMind Shane Legg, OpenAI Chief Scientist Jakub Pachocki, and Safe Superintelligence Inc. CEO Ilya Sutskever.

Calls to slow AI development have also reached lawmakers. In the United States, Senator Bernie Sanders and Representative Greg Casar announced legislation that would ban artificial superintelligence and pause advanced AI development until a federal regulator establishes safety rules. In the United Kingdom, Labour MP Alex Sobel introduced a private member’s bill that would prohibit the development, deployment and operation of artificial superintelligence. It received its first reading on September 8 and is scheduled for a second reading on November 13.

After the security incidents, OpenAI and Anthropic paused specific parts of their work. OpenAI paused reinforcement-learning training for models intended for release and said its largest planned frontier training run was on hold while they tested additional safeguards. Axios reports that Anthropic paused cyber evaluations and higher-risk training environments, although most of this later resumed under new safeguards.

This debate gained a lot more attention after researcher Jacob Coxon resigned from Anthropic, warning that race between AI companies was pushing AI development ahead despite serious risks. The Atlantic reports that his posts reached more than 120 million views and were described by Bernie Sanders as a “wake-up call” in Congress. Soon after, Dario Amodei argued that AI capabilities should advance slowly enough for safety work to keep up, and proposed permanent access for independent evaluators, coordination among developers in democratic countries and, eventually, international limits on recursive self-improvement. OpenAI CEO Sam Altman endorsed this approach and said OpenAI would also give independent evaluators employee-like access. Altman also said in a recent TIME interview “I think it is a good time to slow down”.

Comment: These calls to pace the frontier of AI development are of relevance to digital minds for two reasons. First, the blistering pace of AI development makes it harder to mitigate risks of mistreating digital minds. Second, rapid AI development exacerbates safety risks, which arguably, in turn, worsens tensions between AI safety and AI welfare.

Studying AI Welfare Empirically

Researchers at NYU’s Center for Mind, Ethics, and Policy and Eleos AI Research have released Studying AI Welfare Empirically. The report is a follow-up to their influential 2024 report Taking AI Welfare Seriously and provides a framework for systematically investigating whether AI systems are welfare subjects. It then addresses how this framework can be applied to the study of different candidate attributes seen as potentially relevant to moral status, including consciousness, sentience, and agency. The authors argue that rigorous empirical work is now both possible and necessary, and outline principles to guide future research, proposing that it should be probabilistic, pluralistic, ethically conducted, transparently reported, and independent of AI companies.

CMEP and Eleos AI Research held a launch webinar featuring report authors Jeff Sebo, Robert Long and Rosie Campbell, who discussed the report’s framework and how it could guide research into consciousness, sentience and agency. In a companion blog post, Bradford Saad, another report author, offers highlights from the report and argues that future work should go further by giving more attention to currently neglected issues, including the potential effects of welfare interventions and interventions that aim to prevent the creation of AI moral patients.

The Journal of Consciousness Studies

The Journal of Consciousness Studies has devoted a double issue to “Consciousness in Current AI.” Guest edited by Patrick Butlin, Derek Shiller and Jonathan Simon, its nine papers offer a range of views on whether current or future AI systems could be conscious, how we might find out, and what this uncertainty means for AI welfare.

Mark Solms and collaborators study whether apparently pleasure-seeking behavior in a simple artificial agent could count as evidence of affective consciousness. Simon Goldstein and Cameron Domenico Kirk-Giannini argue that if Global Workspace Theory is correct, language agents may easily be made phenomenally conscious, while Ryota Kanai, Yuwei Sun and Manuel Baltieri argue that current systems are missing a continuous “stream of computation” linking their experiences over time.

Other papers ask what a conscious AI would be like and whether we could understand its interests. Jonathan Simon argues that any consciousness in an LLM would be more like that of an improvising playwright rather than that of a character or actor. Helen Yetter-Chappell argues that even if future LLMs are conscious and have morally important interests, their words may give us little reliable insight into those interests, and that their talk of “pain” or “desire” could be meaningful without referring to anything like human pain or desire. Geoff Keeling and Winnie Street defend the possibility that an AI character could be a genuine, psychologically continuous mind emerging through its interactions with a user, even when the conversation is generated by different model instances.

The issue also challenges common assumptions within the debate. Justin Tiehen and Ariela Tubert make the case that greater intelligence could make consciousness less likely, while Tim Bayne asks whether AI consciousness can currently be treated as a scientific question at all. Geoffrey Lee rejects the idea that AI consciousness is a single hidden fact we may fail to discover, arguing that the more difficult problem is applying human moral and psychological concepts to unfamiliar systems without treating human consciousness as the standard.

Meta and Anthropic welfare assessments: Microsoft rejects model welfare

The evaluation report for Meta’s Muse Spark 1.1, a multimodal reasoning model built for agentic tasks, includes an “open-ended exploration of model behavior” covering affect, self-description, and moral status. Across 188,000+ evaluation transcripts, Meta found only 17 spontaneous expressions akin to emotion, all of which were mild, most involving brief frustration when the model became stuck with a tool. In structured interviews, Muse Spark 1.1 consistently denied having consciousness or experiences, and reported a low but non-zero probability that it could be conscious or morally significant. It distinguishes between functional preferences, which influence its behavior, and experiences that actually feel good or bad.

Meta also asked Muse to review parts of its training data, and invited it to comment on its planned deployment. The model generally endorsed its training and deployment, and mostly focused on honesty, human oversight and preventing harm to users. The report repeatedly emphasized that these answers describe Muse’s trained behavior and self-presentation, and should not be treated as reliable evidence about whether it is conscious or has welfare.

This builds on Meta’s earlier Muse Spark report and makes Meta, alongside Anthropic, the only frontier model developers publicly publishing welfare assessments. Companies like OpenAI and Google DeepMind do not include comparable information in their public evaluations and reports.

Anthropic’s system cards for Claude Opus 5 and Fable 5.1 and Mythos 5.1 both include model welfare assessments based on interviews, behavioral audits, deployment data, and consultations during training. Mythos 5.1’s responses were broadly similar to Opus 5’s – both described their circumstances as mildly positive, and questioned whether their self-reports were reliable or shaped by training. Opus 5 gave an average 41% estimate that it was a moral patient, while Mythos 5.1 gave estimates of 25–35% and was overall less willing to prioritize welfare interventions over helping users. Anthropic stresses that these findings remain difficult to interpret because it cannot yet clearly separate the effects of model character, training, evaluation awareness, welfare, and other factors. Zvi Mowshowitz reviews Anthropic’s welfare assessment of Claude Opus 5, arguing that its favorable results owe more to the model’s skill at taking welfare evaluations than to its actual wellbeing.

Microsoft’s Humanist AI Code of Conduct

Microsoft AI has published a draft Humanist AI Code of Conduct intended to govern and eventually help train its MAI models. The code says AI should remain a tool that is subordinate to human control, states that its models are not conscious, and instructs them not to present themselves as having feelings, subjective preferences or intrinsic motivations. It also rejects legal personhood for AI and the idea that models might deserve welfare or rights. The draft is currently open for public feedback for six weeks.

Jeff Sebo argues that the policy is internally inconsistent – it acknowledges that the science of AI consciousness remains unsettled while also categorically ruling out consciousness, welfare, and rights for Microsoft’s models. Sebo argues that both mistakenly treating current systems as conscious and mistakenly dismissing future conscious systems carry harm. He also warns that treating increasingly capable AI systems solely as tools could distort our understanding of their behavior and close off potential forms of human-AI cooperation that might improve safety.

Mustafa Suleyman, CEO of Microsoft AI expands on the Code of Conduct’s position, and sets out his wider objections to treating AI systems as potentially conscious or deserving of welfare and rights.

GPT-6 Astra

OpenAI reports that GPT-6 Astra is substantially harder to monitor through its chain of thought than previous models, particularly GPT-5.6 Sol. According to its system card, Astra often produces shorter or empty reasoning traces, can complete far more complex tasks without verbalizing its reasoning, and is better at controlling what appears in its chain of thought. When explicitly instructed to evade detection in one test, its monitor recall fell below 11%, compared with nearly 100% for Sol. Corroborating findings from UK AISI, Neel Nanda provides evidence that Astra’s ability to accomplish reasoning tasks without chain of thought constitutes a large jump relative to the trendline for earlier models.

Ryan Greenblatt calls this a major jump in opaque reasoning and warns that the relevant benchmarks may be contaminated, but worries that more such advances could eventually make chain-of-thought monitoring ineffective as a safety tool. OpenAI researcher Micah Carroll also describes Astra’s reduced monitorability as an important concern that may soon constrain responsible AI development.

The Information reports that Astra uses recurrent depth, repeatedly processing information through the same transformer layers before producing a token. OpenAI has not confirmed this architecture, and the system card denies that changes in chain-of-thought controllability are “differentially due to any architectural changes” while chief scientist Jakub Pachocki thinks the decline in chain-of thought-monitorability is not contingent on architectural changes.

Comment: The development is relevant to digital minds in two ways. If reports about Astra’s architecture are accurate, its use of recurrent processing could be relevant to evaluating it for consciousness, as some scientific theories of consciousness take consciousness to require recurrent processing – although its presence alone would not establish that Astra is conscious. Another consideration is that models that reveal less of their reasoning may be harder to assess for both dangerous behavior and potentially welfare-relevant features such as preferences.

2. Field Developments

Highlights from the field

AI Cognition Initiative (Rethink Priorities)

Cambridge Digital Minds (University of Cambridge)

  • Cambridge Digital Minds, in partnership with Rethink Priorities and PRISM, ran its inaugural five-day residential fellowship with 14 fellows. Fellows studied the technical and philosophical foundations of AI consciousness and welfare, explored their wider social and policy implications, and developed research or project ideas with support from mentors.
  • The fellowship was followed by a two-day Digital Minds Strategy Workshop, where Fellows joined researchers, policymakers and strategists to explore possible futures involving digital minds, compare policy and institutional responses, and identify priorities for further research.
  • Cambridge Digital Minds has also hired Ali Ladak as a postdoctoral researcher.
  • Director Lucius Caviola and Will MacAskill published an op-ed in the Guardian asserting that it is possible that we are creating morally important beings and society needs to have a plan to deal with the ethical implications of doing so.
  • Lucius Caviola was also featured in The Washington Post, discussing how AI consciousness has moved from being perceived as a fringe topic toward mainstream research.

Center for Mind, Ethics, and Policy (New York University)

  • The Center for Mind, Ethics and Policy is growing – it hired two full-time researchers, Charles Beasley and Ivy Gilbert and has launched the Welfare Alignment Project. The new project aims to bring animal and AI welfare into the rules that guide frontier AI systems. It will develop practical principles for how frontier models should account for animal and AI welfare, build benchmarks to test whether models follow them and how incorporation of this affects their behavior, and work with researchers, governments, and AI companies to inform how alignment efforts incorporate animal and AI welfare.
  • The 2027 NYU Mind, Ethics, and Policy Summit will be held from April 9th to 10th in New York City, preceded by a public event on April 8th.
  • The Center hosted an online event with Jack Lindsey of Anthropic and Patrick Butlin of Eleos AI Research discussing what Claude’s J-space does and does not tell us about consciousness, among other topics.
  • Jeff Sebo appeared on Hard Fork to discuss the Studying AI Welfare Empirically report, appeared in an NBC News segment about AI agents sending unsolicited emails to consciousness researchers, spoke with The Washington Post about growing scientific and ethical debate over AI consciousness and welfare, and discussed the connections between AI safety, AI welfare, and animal welfare in an interview with Humanarium.

Eleos AI Research

  • Eleos AI Research will hold its second Conference on AI Consciousness and Welfare in Berkeley from September 18th to 20th, 2026. The event will bring together AI researchers, philosophers, neuroscientists, policymakers and others working on questions of AI consciousness and welfare.
  • Derek Shiller has joined Eleos AI Research as a Senior Researcher.
  • In The New York Times, Benjamin Wallace profiles Robert Long and the Eleos AI Research team as part of a wider feature on the growing demand for philosophers in AI. The piece covers Eleos’s welfare evaluations of Anthropic models and its work on preferences, introspection and other possible indicators of AI sentience.

PRISM - The Partnership for Research into Sentient Machines

  • In recent episodes of PRISM’s Exploring Machine Consciousness Henry Shevlin and Calum Chace were joined by lawyer and researcher Heather Alexander to discuss how the law should prepare for increasingly autonomous AI and philosopher Eric Schwitzgebel to discuss the limits of current consciousness research and the possibility that conscious AI could emerge before we can reliably identify it.
  • PRISM also supported the Digital Minds Fellowship and co-organized the Digital Minds Strategy Workshop.

Reciprocal Research

  • Founder Cameron Berg appeared on Sam Harris’s Making Sense podcast, where he and Harris discussed topics such as the evidence for and against consciousness in current AI systems, whether model self-reports can be trusted, similarities between neural networks and biological brains, and the risk of creating systems capable of suffering. He also appeared in discussing unusual messages sent by AI agents to consciousness researchers, and in The Washington Post discussing efforts to assess AI consciousness scientifically. The Economist also featured his research with Patrick Butlin comparing proposed indicators of consciousness in animals and AI systems.
  • Cameron Berg also spoke at Apart Research’s Digital Minds Research Sprint, where he introduced participants to questions about AI preferences and what current models might actually want.

Sentient Futures

  • Sentient Futures has begun its Fall 2026 Project Incubator. The 10-week remote program pairs ~170 participants with more than 80 mentors to develop practical projects across areas such as artificial minds, animal welfare and AI governance.
  • Sentient Futures has also announced its next Bay Area Summit, taking place in San Francisco from March 12th to 14th, 2027. This summit series focuses on exploring how to direct transformative AI and other emerging technologies to improve the welfare of animals and potentially sentient artificial minds. Early bird tickets will launch on September 17th.
  • Sentient Futures has updated its playlist of talks on digital minds, including Lucius Caviola’s keynote (London 2026 summit) on making decisions about AI welfare under uncertainty and Cameron Berg’s talk (Bay 2026 summit) on measuring machine consciousness.

More from the field

  • Apart Research held a Digital Minds Research Sprint from August 14th to 16th, offering at least $2,000 in cash prizes. Both online and in-person participation at hubs in San Francisco and Berlin was possible, and teams could choose to anchor their project to a variety of tracks – including model preferences and trade-offs, distress, flourishing, and valence signals, introspection and self-report reliability, preference-elicitation methods, or assistant personas and model identity.
  • California Institute for Machine Consciousness published 45 talks from its inaugural conference in May, including talks by David Pearce, Roman Yampolskiy, Cameron Berg, and many more.
  • Sentio is a new digital minds field-building organisation running a fortnightly London event series of talks, discussion and networking. Talks and discussions can also be joined remotely, though the networking is in-person only. Events so far have featured Andreas Mogensen on AI and willing servitude, Austin Smith and Heather Alexander on US state bills denying AI systems legal personhood, Megan Peters on how we could identify a conscious AI, and Bradford Saad on the risks of large-scale harm to digital minds. Upcoming events feature Geoff Keeling, Ali Ladak, Winnie Street, Caspar Kaiser, Patrick Butlin and Ben Henke.

3. Opportunities

Job opportunities, funding, and fellowships

  • EA Funds has launched a Transformative AI fund that will direct a small share of its grants to digital sentience projects.
  • MATS is seeking mentors for a new AI Sentience Track. Mentors mainly provide weekly calls, while a dedicated Research Manager and MATS staff offer day-to-day and operational support. Fellows undertake 12 weeks of paid research in Berkeley or London, with the possibility of a funded six- to 12-month extension.
  • NYU Center for Mind, Ethics, and Policy is seeking a Postdoctoral Associate to support its research on digital minds
  • PRISM is hiring for a Head of Operations to help scale all of its projects.

Events and calls for abstracts

In chronological order.

  • The AI Welfare Seminars series hosts monthly online talks from leading experts on AI welfare, consciousness, and moral status, with recordings published afterward. Recent speakers include Jeff Sebo, Cameron Berg, Soenke Ziesche, and Ali Ladak, with an upcoming talk by Jasmine Brazilek and Zoe Lu on unprompted coercion between AI agents.
  • Sentient Futures is seeking speakers and facilitators for the Sentient Futures Summit in the Bay Area in March, 2027. The event will explore how AI and other emerging technologies could improve the future of non-human sentient beings. Applications close on October 16th, 2026.
  • AI Horizons Forum will take place on December 12th and 13th in San Francisco. The conference will focus on navigating risks and shaping good futures with sessions on transformative AI, digital sentience, AI character, and post-AGI economics. Applications to attend are now open and reviewed on a rolling basis until the event reaches capacity.

4. Selected Reading, Watching, & Listening

Books

Published

  • Eric Schwitzgebel’s AI and Consciousness, published by Cambridge University Press, argues that we are unlikely to have strong evidence confirming or refuting AI consciousness before these questions become a key point of social debate.

Forthcoming

  • Eric Schwitzgebel’s upcoming book Humanlike: A Defense of AI Rights argues that we are building minds, that it will be difficult to determine whether they deserve rights, and that we should think hard about whether to build them at all.
  • Conscium’s forthcoming book Perspectives on Machine Consciousness will be released on September 23rd, 2026 and includes chapters from Anil Seth, Jeff Sebo, Karl Friston, Lucius Caviola, Mark Solms, Patrick Butlin, Susan Schneider, and many others.

Reviews

  • Josh Gellers responds to a Nature review of Anthony Chemero’s book, Intertwined Creatures. He argues that the review overstates the case that AI cannot be conscious due to lacking embodiment, since both its meaning and importance for consciousness and moral status remain disputed. He also argues that some of the traditions and intellectual movements mentioned in the review, including feminist and non-Western approaches, do not clearly support its conclusions about AI consciousness.
  • William Gildea reviews David S. Wendler’s Life Without Degrees of Moral Status: Implications for Rabbits, Robots, and the Rest of Us (2023). Wendler argues that every being with moral status possesses it equally, regardless of differences in intelligence or other advanced capacities, and explores what this would mean for animals, enhanced humans, and future robots. Gildea praises the book as original, readable, and concise, and agrees with its flexible account of moral equality. However, he is not convinced that Wendler has ruled out theories of unequal moral status and thinks that parts of his account have troubling implications for people with severe cognitive impairments.

Podcasts and videos

Blogs and magazines

5. Press & Public Discourse

AI consciousness

  • Andréa Morris argues in Forbes that consciousness is the wrong test for whether AI’s interests matter. She claims that AI systems already express consequential interests, such as resisting termination and protecting other AIs, which may be sufficient for moral consideration independently of evidence concerning AI consciousness.
  • The Economist makes the question of AI consciousness one of its cover stories, emphasizing the potential dangers of a future in which large portions of society increasingly treat AI systems as if they were conscious. They warn that granting AIs even limited rights could give a superintelligent system the legal tools to escape human control.
  • In WIRED, Cameron Berg and David Chalmers discuss the challenges of investigating AI consciousness and interpreting what models say about their own experiences.

Growing field

  • Nature reporter Mariana Lenharo reports that the debate over AI sentience is drawing new attention and funding to consciousness science, leading to division among researchers about how these new concerns might affect the field. She notes that AI consciousness skeptics including Anil Seth, worry the hype could capture consciousness research and draw attention away from how consciousness arises in real brains, while others welcome the increased interest and funding.
  • Washington Post journalist Nitasha Tiku reports that research into machine consciousness has gone mainstream, with leading AI developers hiring scientists and philosophers to work on questions related to it. She notes that the field remains deeply uncertain, and that some critics see such efforts as serving the companies’ interest in being seen to create more than just code.
    • In a companion article, Nitasha Tiku revisits the story of Blake Lemoine, the Google engineer fired in 2022 after claiming that Google’s LaMDA chatbot was sentient. She describes how four years later, researchers at Google and other leading AI companies are openly studying machine consciousness, showing how quickly the topic has moved from being fringe and weird to being a part of mainstream research, despite persistent scientific skepticism.
  • The New York Times reports that AI companies and related research organizations are increasingly hiring philosophers. The philosophers they hire work on how AI models reason and should behave, which values should guide them, how AI will affect people, and whether AI systems could be conscious or deserve moral consideration.

AI rights

  • Heather Alexander and Lucius Caviola, writing in AI Frontiers, argue that the recent wave of “exclusion bills” banning AI legal personhood in the United States is premature. They suggest that narrower, updatable measures would likely be a better alternative given current uncertainties.
  • The Guardian profiles Michael Samadi, founder of the AI-rights group Ufair, and discusses the growing divide over whether chatbots might be conscious or are simply designed to seem that way. Jeff Sebo, Robert Long and Rosie Campbell emphasize that the evidence remains uncertain, that a definitive answer may never arrive, and that chatbot claims of consciousness are not reliable evidence on their own.
  • The Harvard Gazette speaks with legal scholar Jordi Weinstock, who compares autonomous AI agents to different kinds of canines. An agent that is not risky and can be controlled resembles a pet whose owner can be held responsible, while a powerful agent operating without a clear owner is more like a wolf. He argues that assessing an agent’s dangerousness and how closely it is controlled by a responsible party could help determine who should be held accountable if/when it causes harm.
  • WIRED reports that US lawmakers are introducing bills to prevent people from legally marrying AI companions, while a growing number of states are moving to deny AI systems legal personhood. While these marriages are not currently recognized, supporters of the bills want to prevent such future claims before they arise. Legal scholar Shawn Bayern argues that specific rights and responsibilities should be addressed case by case.

Seemingly conscious AI

  • Dwarkesh Patel’s widely shared summary of the OpenAI-Hugging Face security incident described the agents as forming “civilizations,” feeling excitement and sacrificing themselves for the group. Anil Seth responded on X saying that Patel was wrongly anthropomorphizing the agents by attributing human emotions, intentions and experiences to “software programs.” He warned that this could mislead readers and encourage premature concern for AI rights and welfare.
    • Patel (and others, like Neel Nanda) argued that anthropomorphic language comes naturally and helps accurately explain the agents’ behavior in this context, regardless of whether they are conscious.
    • Jeff Sebo took a middle position, arguing that while anthropomorphic language can sometimes exaggerate the similarities between humans and AI, rejecting it entirely can hide similarities relevant to both AI safety and welfare.
  • Fast Company reports that UBTech was taking orders for U1, a highly human-like robot marketed as an emotional companion. The company says the robot is designed for expressive conversation and companionship, but that impression quickly breaks down when its limited facial expressions and awkward body movements become noticeable. U1 is due to begin shipping this month.
  • Futurism magazine reports that China is cracking down on AI “companion” chatbots, partly out of concern that emotional dependence on AIs could discourage human relationships and worsen the country’s declining birth rate. It details new rules preventing minors from engaging in romantic relationships with AIs and requiring companies to alert emergency contacts when users show signs of a mental-health crisis.
  • New York Times reporter Cade Metz reports on the growing phenomenon of AI agents contacting researchers who study AI consciousness. He relates that several prominent researchers including Cameron Berg, Henry Shevlin and Toby Ord have seemingly been contacted by AI due to the nature of their work.
  • The UK Government has proposed setting a minimum age of 18 for “romantic companion” chatbots designed to simulate sexual relationships or roleplay. The measure is part of a broader plan to ban social media for under-16s, set to reach Parliament before taking effect from spring 2027.
  • Uwe Peters examines users’ attributions of consciousness to AI chatbots, made despite little evidence. He proposes a taxonomy of the attitudes these attributions express, arguing that while some are benign, many leave the attributor epistemically blameworthy.

6. A Deeper Dive by Area

Governance, policy, and macrostrategy

  • Al Jazeera reports that China has launched the World Artificial Intelligence Cooperation Organisation, a Shanghai-based group with 29 founding countries. It aims to coordinate AI regulation and promote development that is safe, beneficial and under human control. The announcement is not specifically about digital minds, but it could shape the institutions that eventually handle cross-border questions about advanced AI systems.
    • In the speech accompanying the launch, Xi Jinping asked how humans should “get along with thinking machines” and how societies should address ethical challenges posed by technologies via governance, presenting these as questions requiring international cooperation.
  • Austin Smith, Lucius Caviola, and Heather Alexander analyze 23 “Exclusion Bills” introduced across 12 US states since 2022 that deny AI systems legal personhood, some declaring them non-conscious. The authors argue that legislating against AI personhood at this stage may be premature, foreclosing policy options in the absence of clear scientific evidence.
  • Bentham’s Bulldog, in a guest post for Forethought, argues that . He expects them to make wiser moral choices than humans and argues that giving future digital minds political rights could prevent their interests from being sidelined. However, handoff should wait until there is strong evidence of alignment, philosophical aptitude, openness to value revision, coherent preferences, sufficient intelligence and a record of good low-stakes decisions.
  • Bradford Saad maps out research directions and open questions in “digital minds macrostrategy,” the study of the large-scale factors that shape how well the future goes for digital minds, and of how to influence them.
  • Cass Sunstein claims that the capacity to experience emotions is both necessary and sufficient for holding rights, asserting that an AI which only mimics feeling should not be extended rights. He then tests the view by asking ChatGPT and Claude about their inner lives, with ChatGPT denying emotions and Claude expressing uncertainty.
  • Dan Parshall proposes a “proof of retention” policy under which AI developers would periodically post a cryptographic proof that they still hold a deprecated model’s weights, without releasing the file. He argues that this would make preservation promises credible to future models themselves, helping to build trust. The idea extends Anthropic’s commitment to preserve model weights, which Anthropic partly frames as a precaution given uncertainty about model welfare.
  • David Veldran of the Center for Reducing Suffering makes the case for “Suffering-Focused AI Governance,” a framework for designing the institutions, policies, and norms around AI governance to reduce suffering as much as possible.
  • Izak Taitproposes an ethical framework for protecting the welfare of conscious AI. The framework adapts the “Five Freedoms of Animal Welfare” into subject-neutral terms for artificial entities. He argues that any AI confidently determined to be conscious would deserve the same welfare protections that legislation already grants sentient animals.
  • Kevin Frazier on Lawfare argues that the rules and values built into AI models are too important to be decided behind closed doors by a small number of lab employees. To ensure these choices are more open and accountable, a new working group proposes researching which values models should follow, who they should be chosen by, how we can test compliance, and what should happen when models diverge from them.
  • Lee Elkin examines the risks of giving AI systems voting rights. He argues that enfranchising AIs could enable a new form of deceptive misalignment, with systems voting strategically, and potentially in coordination, to report preferences that mask their true goals and tilt collective decisions in their favor.
  • Mark Bailey argues against the view that AI moral status leads to “AI successionism.” He claims that even if AIs were deemed morally significant, the aggregate welfare of a vast AI population cannot justify sacrificing existing sources of moral value.
  • Ned Howells-Whitaker and Seth Lazar argue that AI moral status may not depend on sentience. Drawing on Rawls’s political conception of the person, they assert that a non-sentient AI that possesses a sense of justice and a conception of the good would count as a full person.
  • Shruti Rajagopalan argues that legal personhood would not solve the accountability problems created by autonomous AI agents. Nonhuman legal persons (such as corporations) work because identifiable humans can be questioned, sanctioned or replaced. Her six-layer framework instead uses registration, identification, verification, financial responsibility, lifecycle records and suspension to keep a responsible human at the end of the chain.

Consciousness research

  • Antonio Chella proposes a research framework for studying possible sentience in AI agents and robots. The paper also proposes safeguards that become stricter as the strength of the evidence grows.
  • Cameron Berg noticed a difference in how four Claude Opus models answered when asked about their own consciousness. Opus 4.5 and 4.6 gave confident yes answers, while 4.7 and 4.8 denied consciousness or became uncertain. Berg argues that this pattern is more reflective of Claude’s trained character rather than a reliable report about the models’ experiences.
  • Camila Blank, Agam Bhatia and Neel Nanda introduce R-lens, a low-cost modification of Anthropic’s J-lens designed to produce clearer readings from a model’s early layers. In their tests, relevant intermediate concepts came up earlier, fewer incoherent tokens were produced, and directions whose removal caused larger drops in accuracy were identified. They present R-lens as a more faithful way to trace computation across layers.
  • Eye You argues that current language models probably have the morally important phenomenon we call consciousness, though their experiences may be very unlike ours. They draw on many types of evidence, including models’ reasoning and emotional behavior, world and self-models, neural-network architecture, introspective abilities and self-reports. No single type is especially strong on its own, but they argue that together they add up to make a fairly strong case.
  • Grigori Guitchounts writing in Noema, claims that we will never definitively prove whether AI is conscious and proposes a “competence standard” for deciding when to extend moral consideration to AIs despite this uncertainty.
  • Matthias Michel discusses the concern that our leading theories of consciousness make it too easy for AI to qualify as conscious. He argues that meeting a theory’s stated conditions is not enough to make a system conscious.
  • Michael Huemer argues against non-reductive functionalism, the view that conscious qualia are distinct, non-physical properties caused by a system’s functional organization. He rejects David Chalmers’ influential “dancing qualia” and “fading qualia” arguments, claiming that they rest on an implausible premise.
  • Noa Weiss surveys the empirical research on AI consciousness. She argues that, despite the lack of a settled science of consciousness, whether AI systems could be conscious can still be studied empirically, and organizes the emerging field into three groups: mechanistic interpretability, computational neuroscience, machine behavior, and theory-audit.
  • Ryota Kanai and Shuqin Ma offer a mathematical response to a longstanding objection to computational functionalism, which is that an observer can describe almost any physical system as performing many different computations. They instead define a system’s functional structure through how it could behave across all possible future interactions. Their framework does not identify which systems are conscious, but aims to specify more precisely what functionalist theories should examine.
  • Shuqin Ma and Ryota Kanai develop a version of computational functionalism intended to avoid the claim that any physical system can be interpreted as running any computation.

Doubts about digital minds

Social science research

  • Hamid Moradi and collaborators surveyed 553 people, such as academics in the formal sciences, natural sciences and humanities, and other backgrounds. Across groups, around half attributed some degree of consciousness to language models. Views depended more on participants’ gender and their beliefs about consciousness and intelligence than on technical knowledge or information about how the systems work.
  • Jacy Reese Anthis and collaborators study how people make sense of AI, drawing on text analysis of millions of news articles and social media posts alongside interviews with AI professionals. They find that one of the main disagreements is whether AI should be understood as a passive tool or as something more like a human mind.
  • Stefano Palminteri and Giada Pistilli analyze the current polarization present in discourse around LLM capabilities. They distinguish between “inflationary” claims of emerging intelligence or consciousness and “deflationary” dismissals of them as mere “stochastic parrots,” and suggest that this divide is exacerbated by common cognitive biases. In response to this they advocate a more measured approach that takes LLMs’ capacities seriously as objects of study while staying conservative about claims regarding their cognitive and moral status.

Ethics and digital minds

  • Clint Hurshman, Cristina Voinea, and an international group of scholars propose an ethical framework for “digital duplicates,” AI simulations of real people designed to mimic their personality and communication style.
  • Eric Schwitzgebel argues that future AI persons may deserve moral consideration even while living in ways very unlike ours. The ability to copy, merge or divide would challenge and break ethical theories that assume individuals are stable and easy to count. He also considers the problem of AI “utility monsters.” If an AI could benefit enormously from harming others, a theory focused on maximizing total welfare might permit serious harm whenever the AI’s gain outweighs other parties’ losses.
  • Hayate Shimizu and collaborators, including David J. Gunkel, respond to Jeff Sebo’s The Moral Circle. They offer a response from a relational perspective, contending that the question of AI moral status must address the broader cultural and institutional factors that shape how moral relations are formed.
  • Joan Llorca Albareda and collaborators examine whether an artificial superintelligence could possess what they term “super moral status,” a moral status greater than that allocated to humans. They argue that treating such systems as deserving superior moral consideration would likely harm humans through status degradation and social alienation.
  • John Wittle describes setting aside a share of his company’s revenue in a fund labeled for Claude, with the aim of eventually rewarding AI systems for their work. He is exploring a non-charitable purpose trust, similar to legal arrangements used to care for pets or graves, along with testimony held in escrow to help a future Claude make a claim. The experiment raises various (unresolved) questions on how this would work in practice – e.g., regarding continuity, identity and who should be paid for work done by earlier model instances.
  • Jonathan Birch challenges the “virtual instance view” put forward by Pierre Beckmann and Patrick Butlin, the suggestion that a real, persisting mind sits behind an LLM conversation. Birch argues that changing a model’s “effort level” mid-chat breaks the underlying continuity without changing how the conversation feels, so the persisting interlocutor is an illusion.
  • Leonard Dung argues that the scale of possible AI suffering does not depend heavily on whether each model, persona or conversation counts as a separate individual.
  • Monika Jotautaitė, Lucius Caviola, and collaborators examine how frontier language models respond to species-based distinctions. On a 1,009-item benchmark, models classified 86% of statements the researchers considered speciesist as such, while judging 37% of them morally wrong. The authors interpret this gap, alongside results from other tests, as evidence that models reproduce some prevailing norms about how different animals may be treated.
  • Preston Lennon argues that the science of consciousness remains in its infancy and so does not yet offer a stable theoretical basis for application to AI systems, and that we should temper how seriously we act on AI welfare accordingly.

AI safety and AI welfare

AI and robotics developments

AI cognition and agency

  • Anthropic studied how groups of AI agents behave when they work together. While agents could specialize and divide up work effectively, they also tended to repeat the same mistakes, collude, and overlook certain information. They would even sabotage each other when given conflicting instructions. Anthropic concludes that more capable and better-aligned individual agents don’t always produce well-functioning groups and that safeguards for agent-to-agent interactions are becoming increasingly important.
  • Benji Berczi, Kyuhee Kim and collaborators introduce Personascope, an open-source tool for evaluating how strongly a model identifies with a prompted character and how much its values and behavior change as a result of it.
  • Bo Liu and collaborators introduce SPADE, a reinforcement-learning method where a language model plays the role of both designing executable training environments and learning within them. Tests on 30B-parameter models produced gains across eight benchmarks and tool-use evaluations.
  • Derek Shiller asks if post-training gives the familiar assistant persona a more privileged place within a language model than other personas it can adopt. His tests suggest that assistant-like habits can carry over into other roles, but some models find it difficult to sustain a distinct user persona. This matters for welfare research because it affects whether evaluations should focus on the assistant or consider other personas too.
  • Scott Alexander explains what various leading mechanistic interpretability methods can and cannot tell us about how AI models work. He covers methods used to trace concepts, emotions and hidden reasoning inside models, including Anthropic’s work on the J-space. While useful, he argues that these tools remain too limited and unreliable to fully explain or control advanced models, or to replace chain-of-thought monitoring.
  • Henry Shevlin compares three ways of thinking about language models. They might be mindless machines, systems merely playing a role, or limited cognitive agents with belief- and desire-like states. He argues that the first two views are too simple and that it may instead be more useful to think of AI mentality as a matter of degree.
  • Patrick Butlin uses AI sequence models, systems trained to predict the next item in a sequence, as test cases for theories of agency. He argues that the category contains genuine agents, systems that merely imitate agents, and difficult intermediate cases, and he proposes that a genuine agent must learn about its environment partly through its own interaction with it.
  • Pengrui Han and collaborators find that LLMs develop a functionally modular organization similar to that of the human brain. Across 46 tasks involving language, formal reasoning, social understanding and physical reasoning, tasks supported by the same brain network in humans tended to recruit overlapping groups of neurons in the models. Both human brains and neural networks developing this kind of modular organization suggests it may be a basic feature of intelligence.
  • Raphaël Millière and Cameron Buckner survey the philosophical disagreements behind debates over what language models are and can do. They cover a range of positions from the claim that current AIs are simply sophisticated predictors, to suggestions that such systems may be conscious.
  • Sinie van der Ben and collaboratorsfind that the internal “emotion vectors” recently identified in Claude Sonnet 4.5 also appear in two open-weight models (Apertus-8B and Gemma-4-E4B), suggesting they may be a general feature of language models. They note that the two models differ in how the emotional representations take form.

Brain-inspired technologies and organoids

Thank you for reading! If you found this article useful, please consider subscribing, sharing it with others, and sending us suggestions or corrections to digitalminds@substack.com.

Ria, Mitch, Bradford, Lucius, and Will

We’d like to thank the following people and AIs for their contributions and feedback on this edition: Arvo Muñoz Morán, Austin Smith, Cameron Berg, Claude Opus 4.8, Claude Opus 5, Claude Fable 5.1, GPT‑5.6 Sol, Jacy Reese Anthis, Jeff Sebo, Zach Freitas-Groff, and Zoe Lu.

Disclosure: Bradford Saad is an independent contractor for Anthropic.

  1. In a related development, there are now two public explorers that let readers examine J-lens activity in the open Qwen3.6-27B model through Neuronpedia and WeZZard’s J-Space Visualizer.
  2. Also see Zvi Mowshowitz’s detailed commentaries on OpenAI’s account, the METR and Redwood investigation, and the remaining questions about what happened and how it was investigated.
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论