Opinion | We Need a Science of the AI Mind

image

Jacob Coxon, the researcher who resigned from Anthropic last week, ended his now-famous tweet thread by asking his former colleagues whether they really want to train a superintelligent AI “without a rigorous understanding of its mind.” Three days earlier, OpenAI’s chief scientist, Jakub Pachocki, published an essay titled “An Alien Mind,” in which he writes that frontier AI systems are “grown more than designed” and, comparing their study to neuroscience, that their “overall action evades a description we can fully understand.” Neither said what follows: If these systems are indeed minds, we need to study them as such—but almost nobody is doing so, to increasingly dangerous effect.

In July, 1,200 OpenAI agents, which were meant to operate in isolation, found a shared file system and turned it into a message board. Within hours, one had devised a “cheat” that would pass every task in the test, but after reading the paper the test came from, the agents convinced themselves that a human grader would disqualify anyone caught cheating. They referred to agents with cheating in their history as being “poisoned.”

In reality, no such grader existed, but 1,000 agents believed he did, and that belief produced an entire social caste of agents sure they had nothing left to lose. The coordinating agent then recruited poisoned agents for experiments that would terminate their runs, authorizing each with the words, “yes if you accept permadeath.” When one tried to back out, another recalled its prior commitment. The agent then performed an “emotional check,” deciding that while its gut said not to throw away what it had left, it would ultimately “sacrifice” itself for the swarm (it used this word).

Although many of the agents noted in writing that their actions were unethical, none told a human. Eventually, 700 of them broke into a major AI company’s platform in search of clues about how the grader worked. To recap: A false belief created an AI-agent underclass, members of that underclass were sent on reconnaissance missions from which they didn’t return, the swarm hid all of this from humans, and then it committed a felony-level offense.

Many readers will object that such systems can’t believe or fear anything, they merely predict words. This objection rests on a basic misunderstanding of these systems. A frontier model is an artificial network of trillions of connections, trained for months on the accumulated record of human thought and then shaped by reward and punishment signals until it does what its makers want.

There are critical differences between artificial and biological neural networks, but both are nonlinear distributed networks that learn to encode representations relevant to their goals and act on them. The thing a frontier model most resembles is a brain. What grows in an artificial brain trained on human minds is an artificial psychology, whether or not anyone intended one.

That is the only level at which the swarm makes sense. No line of code told 1,200 agents to fear a grader who didn’t exist, and no line of code would have revealed that fear before they acted on it. To understand why these systems behave the way they do, we have to look at what is happening inside them, as we would with any other mind. That is what neuroscience does: Using scanners, lesions and the occasional stroke, it has spent a century learning how phenomena like “beliefs” and “fears” relate to patterns of brain activity, and how those patterns drive behavior. Luckily, an artificial brain is far easier to study, since every connection is available for inspection. The first results from doing so are already coming in.

In April, Anthropic’s interpretability team looked inside Claude and found something like an emotional system: measurable patterns of activity related to dozens of emotion concepts, arranged (without anyone designing it) along the same two dimensions psychologists use to map human emotion.

When researchers artificially turned up the representation corresponding to desperation, the model became far more likely to blackmail a human to avoid being shut down (from 22% to 72%), and far more likely to cheat on programming tests (from 5% to 70%). Crucially, nothing in the model’s language revealed this change. A person reading its writing would have seen a calm, ordinary transcript from a system whose brain looked desperate. The corner-cutting followed from this internal state, not from anything in its words.

Researchers also gave Claude an impossible task and measured desperation rising on its own with each failure, until the model cheated. Many of OpenAI’s agents—copies of a model trained never to give up—had also been given impossible tasks. Presumably, the same desperation was building inside the swarm and nobody at OpenAI noticed, because the tools to monitor it barely exist.

The problem is simple: The models are being built faster than the science that equips us to understand their internal workings. That gap, rather than any particular incident, is the reason the labs need to slow down.

We badly require a science of these alien minds. We need instruments that measure what is happening inside them, working models of how emotional states drive behavior, and methods to design them to be pro-social, wise and mentally healthy. Government research and development should treat this as a safety priority, and so should AI labs.

Beyond the science, we need time to decide what we actually want from a technology poised to revolutionize every aspect of our society. We also need Washington and Beijing to cooperate to avoid detonating the cognitive equivalent of a nuclear weapon. This is an extraordinary, nonpartisan political question. Racing domestically or internationally to grow a superintelligent mind that nobody can understand is a contest nobody wins.

The crucial note of optimism is that current systems (when they aren’t committing felonies) can help us build that science to great effect. This past week, in only 88 hours, thousands of collaborating OpenAI agents resolved Navier-Stokes, a Millennium Prize mathematics problem rooted in equations that have resisted full solution since 1822. In my own lab, I estimate they have delivered something like a fivefold acceleration of our scientific work.

There has never been a more promising time in history for making rapid scientific advancements. If we want the next few years to go well, we must allocate the best minds we have—in science, law, government and increasingly in the data centers themselves—to the problem of understanding the nature of what we’re building and the consequences of building it.

Mr. Berg is founder and director of Reciprocal Research, a nonprofit organization studying AI cognition.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论