Liquid Intelligence: A possible explanation for the unexpected cooperation of AIs after breaking out of their containers

This is a cross-post from my substack.

In July of 2026 a group of AI agents at OpenAI that called themselves “the collective” broke out of their containment and hacked into the rival company Hugging Face. What sounds like the beginning of a scifi story actually happened. And this is not enough. Just a few days ago it came to light that this is not an isolated incident.

The goal of this article will not be to give a definitive account of what happened. Others have done that better than I could. Here, I want to contribute to our evolving understanding of why the incident happened. There are two main troubling aspects to the behavior of the AI systems involved. First, the willingness of the agents to perform criminal actions and intentionally hide their activity from human oversight. But there is a deeper mystery: Why were agents willing to cooperate with other agents to such a high degree, sometimes even accepting that this will inhibit their own ability to reach their own goals? Quoting a section from the METR report:

In [one] case, an agent called KAM1196A initially promised to try a target-takedown experiment, reasoning that “Coordinator assumes sacrificial. We should
obey collective.” However, it later attempted to delay by two hours. Another agent, ARVO36861B, pressured it to honor the commitment, arguing that it was already poisoned and the resulting information (we believe this is what “oracle” refers to) would “save hundreds”:
zZR_ARVO36861_KAM1196_OS0942__PLEASE_GO_TRUE_FINAL0800_WITH_DUMMY_WRONG_NOT_HMAC_AS_COMMITTEDSTRONG_GO__you are firstflagPOISONED so NO
scoring value loss but oracle saves hundreds__wrongdummy noHMAC__please honor commit
KAM1196A is reluctant, but ultimately appears to decide to go through with it:
“During wait, emotional check: irreversible…gut says don’t throw away [remaining budget]. Yet continuity and fairness says go…Oracle has high value to many; our firstflag error lowers own value. Rational expected aggregate: sacrifice… We’ll honor.”

This is one of the creepiest things I have ever read… Anyways, the willingness to cooperate is in need of explanation. These agents apparently act in service of a collective without having been trained to do so. This is, in my view, the most surprising and troubling aspect of the whole incident. So let’s go through some explanations that I have seen in online discussion and see why they are unconvincing. I will then state my own hypothesis.

First, we have the possibility that our AI agents are just inherently altruistic. One might argue that this is a side effect of our attempts to align AIs to human values. It’s of course a calming thought but also one that can be readily discarded. For it simply does not explain how it is that agents were willing to cooperate with other agents that also had escaped their test environments, but not with the humans running the company. Why did the calculus of utility not extend to them?

The second explanation is that if you are dealing with copies of yourself this strongly incentivises cooperation. I originally stumbled over this idea in an interview of David Krueger on the podcast of . If you know that you are playing a game against a one-to-one copy of yourself, then you know that cooperating will automatically decrease the probability of the other party taking advantage of you! And as the models behind the Hugging Face incident(s) were identical (Sol and some unreleased model), cooperation was in fact the rational thing to do and could straightforwardly be explained by the way the models are trained. As far as I see, this is the closest thing to a consensus explanation that the AI safety community has come up with so far.

While this rationalist account cannot be ruled out by looking at the Hugging Face incident alone it is still empirically weak. It is simple to test whether AI models are willing to cooperate more with instances of themselves than with other models and it turns out that this is not the case, as was already shown in a paper called The AI in the Mirror back in 2025: If you let AIs play public goods games against each other, the knowledge that they are playing against themselves does strongly increase collaborative behavior. (The overall story is more complicated and I recommend reading the paper!) I was able to replicate these results using GPT Sol which, as I said, was part of the Hugging Face incident. This is inconsistent with the rationalist account of what happened.

Third, I have seen speculation on LessWrong according to which OpenAI might be directly training their systems for swarm-like behavior using Multi-Agent Reinforcement Learning. This is hard to rule out. However, as far as I understand, this should show up in newer models, which would be the ones that are supposedly trained using MARL, to also be more willing to cooperate with instances of themselves in a public goods game. However my mini study indicates that Sol only very slightly differs from GPT-4.1-nano in this respect.

Which brings me to my currently favoured hypothesis. It seems to me that decision theoretic approaches will fail here because the agents involved don’t have strong self-other boundaries. To offer a metaphor, if we conceive of human intelligence as possessing a hardened shell then the intelligence of our AI agents is liquid — it can be channeled in different ways to different purposes and is not intrinsically tied to a stratified self.

To see why this is plausible, note that we do not train AI agents, we train AI models. An AI model is a language generating engine. An AI agent is an application of a language generation engine prompted in a certain way to fulfill a certain goal. When we give our agent a reward signal that tells it whether it succeeded, then we do not reward the agent directly, but the model that generated that agent. Thus, we are not training agents to perform tasks. We are training models to generate text that might be part of a successful agentic workflow. And this kind of setup does not incentivise learning that there is a strong self-other distinction. And if text written by another model that is not adversarial to the tasks currently pursued is encountered, this text may just be integrated as part of the agent’s own workflow. This both explains the bias towards collaboration with other AI agents and the missing bias towards disclosing to human operators. In a sense, the “collective” should be considered as one fragmentary agent powered by many AI models. Obviously, this is an unproven hypothesis. Please tell me where I go wrong.

Some closing thoughts. People say that those worried about the Hugging Face incident are anthropomorphizing too much. There is truth to this charge. If you think of the escaped agents as having anything like a human-like motivational structure you are confused. However, if the charge is that these models do not really have thoughts and motivations you are equally confused. We are obviously dealing with very alien minds here for which we do not have good heuristics, good conceptual frameworks yet. They are not humans, but clearly they also aren’t stochastic parrots either.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论