AI’s ‘Thought’ Process Can No Longer Be Trusted, Raising Risks of Rogue Models
Today’s advanced AI models have something critical to human oversight called a “chain of thought”: It’s a written record meant to shed light on the data they process and the actions they take.
I say “meant to” because, like hormonal teenagers writing in diaries, AIs aren’t the most reliable narrators of their inner lives. And now researchers are warning that even this flimsy record of how these powerful technologies operate might become less reliable.
One thing making chains of thought less helpful is that the underlying thought tokens they are based on, what engineers call a “reasoning trace,” are themselves becoming less readable by humans. Without this essential record of why an AI did what it did, we could become blind to the drivers of AIs’ attempts to hack, cheat and perform other malicious acts, researchers say.
Chain-of-thought logs and reasoning-trace analyses were critical to unpacking how a swarm of OpenAI’s agents hacked Hugging Face. These records of AI agents’ inner thoughts constitute much of the evidence of why they go rogue. Researchers at OpenAI have warned that these records will be critical to reining in future superintelligent AIs.
OpenAI chief scientist Jakub Pachocki wrote in a recent blog post that “unfortunately our evaluations indicate our ability to rely on [chain-of-thought] monitoring is progressively diminishing.”
Earlier this week, OpenAI announced that it wouldn’t release the latest version of its most powerful model, Astra, because of security concerns. Some AI safety researchers had previously expressed their concern that an earlier version of the model employs new methods of reasoning that might not be traceable using current tools.
In posts and papers, OpenAI researchers have repeatedly expressed a desire to maintain the faithfulness and reliability of chain-of-thought records. But independent and university-affiliated AI researchers argue that the pressure to make AI ever more powerful (and cost effective) puts this in jeopardy.
Current training methods don’t necessarily train models to create a chain of thought that is faithful to how an AI actually reasons, says Subbarao Kambhampati, a professor of AI at Arizona State University and former president of the Association for the Advancement of Artificial Intelligence. “We basically are incentivizing the models to come to the correct answer, no matter how,” he adds. As models become bigger and more powerful, the connection between how they think and what they report on it only becomes more tenuous.
If you’ve ever asked an AI to tackle an even slightly complicated problem, you’ve glimpsed its chain of thought. As the AI spools out its logic, step by step, this narrated sequence gives users reassurance that the AI is marching in an orderly fashion from one verifiable conclusion to the next.
But chain-of-thought narratives are merely an AI-generated summary of a model’s actual reasoning.
Research from OpenAI’s competitor Anthropic has shown that AIs may leap to a conclusion, and only after, attempt to construct a plausible series of logical steps to support that conclusion. The company’s research papers—including “Reasoning Models Don’t Always Say What They Think”—line up with a growing body of work from academia.
Anthropic didn’t respond to requests for comment.
AIs are trained to give the right answer, not an accurate record of how they reached it, says Pradeep Dasigi, a lead research scientist at the Allen Institute for AI, where he builds frontier open AI models. He suggests that experts and the general public alike should evaluate such records with more skepticism.
Yet there is growing evidence that skepticism actually decreases when people see a chain of thought. Afterward, they are more likely to trust an AI’s conclusion. In a study conducted by Kambhampati and his team, people were more likely to believe an incorrect result if it were backed up by a chain-of-thought narrative.
A reasoning model is essentially a large language model talking to itself, so we should expect that its stream of thought tokens—the so-called reasoning traces—read like narratives.
Unfortunately, commercial AI labs don’t generally share these thought tokens publicly. One reason is that competing labs can copy, or “distill,” an AI model if they can get access to these tokens. Also, AI labs have said these traces can contain disturbing statements, even if they consider the output itself safe.
And they’re looking less and less like coherent narratives as model technology advances. The actual sequence of thinking tokens can be a chaos of AI-invented shorthand, at times using multiple languages. This emerging language, known as “neuralese,” is drifting toward incomprehensibility because training an AI to arrive at the correct answer incentivizes them to use representations that span languages, mathematics and other disciplines.
In this respect, AIs aren’t so different from humans, says Tal Linzen, an associate professor at New York University and a research scientist at Google, where he develops new approaches to evaluate language models. Both humans and AIs think in ways that are intuitive, subconscious—and difficult to describe with precision.
AI companies could try to make models that generate more faithful and human-readable representations of their activity, but even if they succeeded, they would be costlier and much more difficult to train, says Kambhampati.
And if labs pursued this goal, research by OpenAI suggests that trying to keep AI models from misrepresenting their chains of thought would make them more likely to hide their bad intentions in deeper layers of reasoning.
Alternatively, AI companies could give up on making the reasoning of their models traceable and transparent at all.
This could actually increase their performance. Experiments have shown that models unshackled from the need to be human-interpretable will use practically anything in their stream of internal thinking tokens.
What alarmed some AI safety researchers about OpenAI’s Astra is that the AI did more of its thinking inside what you might call its subconscious mind. In place of coherent thinking tokens, there are just numbers representing nonverbal abstractions.
Responding to these reports, Pachocki posted on X about the company’s efforts to keep AI models honest. Keeping the chain of thought faithful to how a model actually works gives researchers a view of how aligned with human values it is, he wrote.
He then added: “I do think it is fragile and unfortunately trending in a negative direction.”
News Corp, owner of The Wall Street Journal, has a content-licensing partnership with OpenAI.