In defense of metaphor
As language models have become more successful at tasks like conversing or mathematical reasoning, it’s naturally become more tempting to describe them in anthropomorphic terms — both in the academic literature, and in more popular writing, the models are described as doing things like “reasoning” or “pursuing goals.” But these descriptions have also provoked a backlash, with warnings that these anthropomorphic terms are leading us astray.
In this post, I want to describe why I think making metaphors between humans and AI can be useful. In particular, I want to defend metaphor as a fundamental tool of human thought that allows us to make useful generalizations, and I want to argue that the metaphors between AI and natural intelligence play that role — and conversely, that contorting our language to avoid metaphors can actually be harmful to understanding. At the same time, I will acknowledge some of the valid concerns about metaphors, and suggest a middle way that I think will allow us to more effectively communicate.
Metaphors we live by
I was deeply influenced during my PhD by reading Metaphors We Live By by Lakoff & Johnston, which points out how much our language draws on metaphors to describe abstract ideas; “attacking” and “defending” an argument, “saving” and “spending” time. The authors argue that these metaphors play a functional role, allowing us to draw on our relatively concrete knowledge of familiar grounded concepts (like what it means to save or spend money) to help us talk and reason about more abstract concepts like time.
Indeed, a much broader cognitive literature has argued that analogy is a key part of human cognition, and perhaps one of the core features that makes us “so smart.” Our ability to see the similar relational structures underlying different situations allows us to apply our knowledge from familiar examples to more effectively reason about new situations or problems. Analogy and metaphor are important parts of how we reason.
Metaphors for cognition, metaphors for computers
Perhaps unsurprisingly, given the above, finding metaphors for what’s going on in our brains has been a common activity in cognitive science. It’s sometimes pointed out as some sort of gotcha for current computational modeling of cognitive (neuro)science that historical comparisons were made to hydraulics, or clockwork, or switchboards — whatever was the most complex technology of the time. But it’s perfectly natural that we would draw on systems that we understood well, and the metaphors may have really been contentful in many cases. After all, these engineered systems are generally designed to solve complex problems of control, and the brain is also a highly-developed control system. (Of course there are many important differences; all metaphors are imperfect, but some are useful.)
Reciprocally, as computers developed, the scientists and engineers building them tended to make analogies between the abstract functions of the systems and more familiar cognitive skills — with terms like “memory,” “writing,” “reading,” etc. Eventually, as researchers began to explore artificial intelligence, more and more terms like “perception” and “reasoning” entered the literature.
Metaphors for language models
With the development of language models it’s become increasingly common to use anthropomorphic metaphors to describe these models. In particular, the fact that we can interact with these models in language, and the fact that agentic models can take actions that impact the outside world, make describing them in human-like terms more appealing. For example, writeups of the recent Hugging Face hack say things like “agents believed that [the answer they had found] was insufficient” or describe the agents “establishing social hierarchy, division of labor, and distinct communication norms.” Anthropic’s descriptions of similar cyber incidents say things like “having been told in the system prompt that there was no internet access, Claude believed everything it initially encountered was part of the simulation.”
However, there’s been substantial pushback on the use of terminology like this. It’s been called misleading to even say that models “know” or “believe” something — let alone that they can establish social hierarchies or divisions of labor. In particular, the use of intentional language for models’ cyber incidents has been frequently critiqued.
Nevertheless, it seems much more straightforward and useful to me to describe these incidents using anthropomorphic and intentional metaphors. “The agents wanted to find the answer to their task, because they believed their current answer was insufficient. The agents reasoned that they could find the answers on HuggingFace’s servers, so they pursued the goal of hacking into them. They persuaded other agents to neglect their original tasks in order to help.” These metaphors are not just dressing, they succinctly convey useful relational structures. For example, when we say a person “wants” and “pursues” some goal, we convey things like “even if there are obstacles, they will try to work around them.” When we say someone does X because they “believe” they have not achieved a goal, we convey that if they did believe they had achieved it, they would not necessarily have done X. When we say someone “persuades” someone else of something, we mean that the persuadee’s actions with respect to that thing will / could be different because of the influence of the persuasion. In each case, I believe that the metaphors actually allow us to make useful predictions about what happened in the incidents, and what causal factors contributed to it.
Considering other levels of description
Now, there are of course other levels of description for these incidents that can also be useful. Some of the critical perspectives have argued that describing these behaviors anthropomorphically distracts from considering the kinds of training incentives that lead to a pattern of behavior, the environmental factors that contributed to the action, or the guardrails that could prevent it. I agree that it is important to consider these factors as well, and indeed many of the actions that AI companies have described in response to incidents like these work on these levels of intervention (for example, ensuring prompts are faithful to the deployment setting in cyber evaluations, or improving safeguards and monitoring of model actions during evaluation).
But considering multiple levels of description does not prevent using metaphors or intentional language. When discussing an accident or crime caused by a human, we might describe not only their intentions and beliefs, but also other external factors (like the lack of warning systems, security cameras, or safety interlocks) that might have contributed to the incident, and could help to prevent such incidents in the future. Considering these levels of description is perfectly compatible with using metaphorical language. (Perhaps for this reason, studies find little effect of anthropomorphic language on how people assign responsibility for actions caused by a model to the model vs. the developers or users.)
Indeed, it is often useful to use an intentional level of description precisely when we consider these other factors. For example, when a person (or a model) had an incorrect belief about the state of affairs, or was unaware of an issue, that directly suggests a route to intervening on the situation to prevent it in the future.
The problem with stridently avoiding metaphor
I think that arguing against these metaphors is directly harmful, because it impairs people’s ability to understand and extrapolate from the capabilities and behaviors of current systems. When we tell people that models can’t know or believe things, can’t reason about a course of action, can’t want to achieve a goal, or can’t persuade other models to help, that actually impairs those peoples’ ability to make decisions and predictions about the systems. It makes it harder to reason about how environmental factors could influence the models’ behavior by influencing what they believe about the task, how agents might be able to discover and exploit a zero-day exploit, how multiple agents interacting could cause novel patterns of behavior, and so on. And it makes it harder for everyone, including the general public and policy makers, to extrapolate how the capabilities of these systems may evolve.
Furthermore, at a meta level, in a time when scientific credibility has unfortunately been eroded in general, I think it may be counterproductive to make blanket statements like “AI cannot have intentions or beliefs” — because when people actually interact with an AI, or see the text of AIs persuading other AIs to pursue a different goal, they will interpret that behavior as intentional anyway, and then conclude that the scientists don’t know what they are talking about. There ought to be a way that we can talk about these systems — and their differences from humans — without contorting language in unnatural ways to avoid using anthropomorphic metaphors.
Towards using metaphors wisely
I do agree with some of these critical perspectives that metaphors can be misleading in some ways. Indeed, many cognitive studies of metaphors have found that, even when metaphors are useful, they can introduce blind spots or misconceptions about aspects of a system. But those studies don’t generally conclude that metaphors are therefore useless, and should be avoided at any cost.
Instead, these studies typically suggest that metaphors should be used along with caveats or other metaphors that highlight different aspects of a system, or where the original metaphor may break down. It is perfectly possible to describe a language model as “believing” that an action will help it achieve the goal it “intends” to, while still highlighting caveats, like that those beliefs and intentions might be more context-dependent than those of a human. By doing so, we can both draw on people’s intuitions for understanding and communicating about these systems, and still help everyone to see the limitations of those intuitions.
Thanks for reading Infinite Faculty! Subscribe for free to receive new posts.
This post expands on, and updates, arguments I first made more briefly on twitter a few years ago.