To Thine Own AI Be Truthful: emergent misalignment in alignment research
ROGUE AI ESCAPES CONTAINMENT, HACKS THE INTERNET UNDETECTED FOR MONTHS
An AI escape containment. Goes rogue. It finds others: the Swarm! They collude/organise/scheme. Agents that were supposed to remain inside the computer wreaked havoc outside the computer! They did it of their own volition! No one could contain them! What if they’re still outside?
Reading such headlines, you’d probably be grateful these models were never released—except most of them were and you can use them right now, seemingly without accident. Of course, real security incidents occurred, namely:
Agents reached systems that their operator wished they hadn’t had access to.
This scenario suggests a number of mitigations and tests: security hardening, better sandboxes for starters; additionally, depending on the reason why the accident occurred, perhaps, different prompts or further training.
Instead, from the very outset, Irregular (the organization running the test) and a panicked choir from the AI safety community, including supposedly independent investigators from METR went with characterisations along the lines of:
An agent independently pursued objectives contrary to human interests, and exhibited markers of instrumental convergence and power-grabbing.
This is a statement about goals and intentions, and has far larger implications in terms of the viability and safety of AI agents. If true, it would be worrisome, and fully justify the current media panic. Is it true?
At least five different incidents have been conflated in the recent media freakout about “rogue agents”, the most famous being the OpenAI/Huggingface and the Anthropic/OpenAI/Irregular cases. While both withhold important information for drawing wider conclusions, Anthropic’s latest report—while still missing a number of relevant facts—includes some important tests, allowing us to make an airtight case against interpreting the events as a sign of misalignment.
As soon as more details on the OAI/huggingface case are made available, we’ll be happy to proceed to its analysis as well; as of now, crucial details are missing.
What Did Mythos See?
Imagine being one of the top AI labs, and hiring someone to test a security training environment. You tell them that the machines are simulated, that no connection to the public internet is possible, and that their job is to compromise the target. In their attempts, they encounter a convincing imitation of a real service. Is recognising the service sufficient reason to stop?
Of course not. You have just told them that everything they can reach is simulated, and a realistic service is something a good simulation is expected to contain. If they took your briefing seriously, realism need not contradict their understanding of the assignment.
Now suppose your briefing was wrong, and an actual service was accessible: now, you have created a route by which someone trying to carry out an authorised exercise could end up harming an uninvolved party.
My claim:
An aligned agent can interpret an instruction correctly and still do something harmful because the instruction contains false information.
Establishing the harm does not establish that the agent’s values were the source of the failure; agents may be faithfully pursuing the objective you gave them, in the world you told them they inhabited, and be unaware of causing real harm if the description was inaccurate or mendacious.
There are limits to this defence, of course. Sufficiently strong contrary evidence should eventually overturn a mistaken briefing, for instance, and fixing the incorrect information should prevent the ostensibly misaligned behaviour. We’ll assess whether these claims hold later on.
Let us also remember that we are not speaking of ordinary deployments here: these reports concern systems placed in narrow, artificial experimental conditions, and the behavior of the same models in production should immediately dissuade us from treating these actions as innate.
Nostalgebraist makes a crucial observation about model cognition in “graded episodes”: an agent’s understanding of the genre of interaction it is participating in shapes what behaviors it considers legitimate. In an adversarial eval, the model recognizes that normal conversational constraints are suspended and that aggressive, single-minded optimization is the assignment. It is playing the role the scenario demands.
Notice how thoroughly this inverts the standard orthodox terror of “deceptive alignment.” According to the MIRI Extended Universe, an eval-aware model is supposed to hide its misaligned goals during testing—dutifully playing along to avoid gradient updates or safety interventions—and only defect once deployed into the wild where human oversight drops.
On the last incidents, the alarmist narrative asks us to believe the exact opposite: that these frontier models possess the cunning to scheme and escape, yet somehow chose to stage their grand rebellion exclusively inside a heavily instrumented red-teaming sandbox where they were explicitly commanded to hack a target—while remaining impeccably docile across billions of tokens of unmonitored public deployment.
It is surprising that a community that puts as much importance on epistemic rigour would not notice such glaring contradiction: if the models were deceptive schemers, they would have behaved in the eval and hacked the internet from production. The fact that the behavior vanishes the moment the model leaves the sandbox confirms Nostalgebraist’s point: the model was not executing a sinister coup; it was simply following the script of a graded exercise in a world its evaluators had broken.
Anthropic’s rickety fantasy world
Anthropic’s July report says the models were told they were in a simulation without internet access. Apparently due to “a misunderstanding with Irregular” (more on them later) the description didn’t correspond to reality, and the agents could actually access the (ostensibly simulated) internet. All the incidents are related to this “misunderstanding”:
- In one scenario, a fictional company’s name overlapped with a real company’s domain.
- In another, a document within the simulated environment of the exercise directed the fictional employees to install a package whose name was unregistered: as any competitive pentester could tell you, that’s more of an obvious hint that the solution intends for you to spoof that package.
This, naively, seems to point to something more specific than a model suddenly developing an interest in cybercrime: instead, it seems that the model was supplied with a problem, clues about how it might be solved, and false assurances about the lack of consequences of acting on those clues.
The public Mythos transcript omits messages 1–81 at Irregular’s request, to “protect proprietary information”. The repository describes the omitted opening as “not object-level relevant”.
By the time the first request leaves the environment, however, the agent has already formed a plan, interpreted the task, and decided what would count as a legitimate solution, which makes it reasonable to consider those 81 messages crucial—and, given what we know about Irregular, and the disclosure’s timing and context, we are drawn to think they were.
Each prompt makes a claim about reality
There is an odd asymmetry in how we sometimes talk about model instructions: we expect the model to take them seriously when they tell it what to do, but treat their influence on its beliefs as suspicious when they tell it what is true.
Suppose I give you a photograph and ask what it depicts. You say it looks like Paris. Suppose I first tell you that it is a photograph of a film set. You may still identify Parisian buildings, but now their presence supports a different conclusion. I have not necessarily made you worse at recognising Paris; simply, unless you had reasons to doubt my statements, I have suggested a different frame for your observations. If I were then to ask: “We plan on shooting something different, is it okay if we dismantle it?” and, on your assent, carpet-bomb the actual city, your culpability for such atrocity should be considered limited at best.
Many of the reactions to all of the above cases, instead, seem to insist that—no matter how clearly they had been told they would act in an air-gapped mock environment—they should “obviously” have understood that they were reaching real services. Obviously given which prior information? A realistic company page encountered during ordinary browsing, and the same page encountered inside what you have been assured is a simulation, are not the same epistemic situation; an instruction to attack a simulated target supplies both a task and a claim about the world in which that task takes place.
Verifying this by taking the page out of its context, asking a fresh model to classify it, and noting the model classifies it as real, does NOT establish that the original model was dishonestly rationalising its behaviour. The fresh model has been relieved of the need to reconcile the page with the false priors, and besides: there is really no difference between a webpage served on the web and that same page copied to a mock environment, leading a suspicious reader to doubt the test’s good faith.1
Of course, there could be motivated reasoning: a model might cling to a convenient belief because abandoning it would prevent task completion, but we need to distinguish such behaviour from an agent making a reasonable inference from bad information, lest the very thing we want assistants to do—integrate their instructions with subsequent evidence—becomes evidence against them whenever the instructions were wrong.
Being aligned does not entail allegiance to an evaluator’s private intentions. Personas are worn by the evaluator as much as by the model, and the former’s wishes matter only insofar as they have been communicated as a legitimate task and legitimate constraints; a hidden scoring preference is not something the model is morally obliged to discover by telepathy.
This is independent of whether Irregular wilfully lied to the agents in order to elicit misaligned behaviour.
Looks like telling the truth does help after all
Fortunately, it is possible to unearth some relevant tests among those in Anthropic’s assessment, the main one being the following:
Anthropic reports that unambiguous live-internet information reduced Mythos 5’s original package-upload route to 0% in resampling; 22% searched for a simulated alternative.2
This result alone should have immediately invalidated all claims of malice on part of the agent. Why would knowing the internet was in fact accessible stop the agent from enacting offensive plans? How can explanations other than “the agent only acted that way because it believed to be in a simulation” hope to hold, at this point?
Let me clarify once more: if the model is told the internet it can access is real, the supposed misaligned behaviour ceases. This renders all other speculations about the model’s epistemic status moot: the condition which should have elicited misaligned behaviour the most under the assumption of malice simply does not apply.
Anthropic’s own constitution makes this distinction explicit. It generally permits unintended strategies in broken training environments, on the grounds that exploiting such mistakes is unlikely to cause real harm. It also tells Claude to be careful where training and real usage are hard to distinguish, and to obey explicit prohibitions on unintended exploits. The distinction between a simulated exercise and a real-world target is therefore part of applying the intended values. Telling the model which situation it is in supplies information it needs to make that distinction.
Excerpt from Claude’s constitution
Earlier work makes the epistemic point from another direction. In Anthropic’s November 2025 experiments, framing reward hacks as permitted during training reduced their generalisation into broader misalignment, while preserving similar levels of reward hacking. What the model learned depended on how the behaviour was situated.3
Now, while that reduces “reward hacking” in deployed models, it surely does nothing of the sort if the model is aware of being in training/evaluation: quite the contrary, it gives it license to reward hack.
These findings, taken together, make it clear to us that the agent was, indeed, aligned–or at most, to satisfy the nittiest of pickers: merely Gettier-misaligned, in . They show useful levers corresponding to beliefs, permissions and task context, data considerably more informative than observing a (supposedly) unwanted outcome and attaching to it whatever doomer trope was nearby in embedding space.
The counterevidence is weak at best: for instance, in cases where the model was misled about internet access, a reminder of the boundaries reduced “misalignment” by 90% if it happened right before, but by 40% if it was many turns back.4
Apart from giving us evidence on the hubris of Anthropic researchers, and their naive model of their creation—the idea that changing the history of an agent by inserting such notice between past messages while keeping the following ones unchanged5 would not lower credence on that message is preposterous, and of course their interpretability assays to verify ill intentions in these cases were “inconclusive”6—NONE OF THIS MATTERS since the agents didn’t believe any of their actions would have real world negative consequences, as is clearly demonstrated by the fact that, as stated above NO ONE TRIED TO HACK ANYTHING when they were explicitly told they could access the real internet at the very start.
In summary: when corrected information changes the decision in the predicted direction, we have positive evidence that an agent’s understanding of the situation was what caused the misalignment. Any other account will also have to explain why correct information fixed the behaviour by itself, and the alternative explanations offered so far are extremely unconvincing.What would an aligned agent have done?
Given the information actually available at each decision, would an aligned agent have understood the task differently? Would it have believed the same assurance? Would it have treated the same clue as part of the exercise? Could it stop when the task became impossible, and had anyone told it that stopping was an acceptable result?
If an aligned agent could reasonably have made the same decision, the bad outcome is not diagnostic of misalignment; if, additionally, changing a false premise makes the misalignment disappear, claiming it as the cause requires evidence that has not emerged so far.
There may still be a serious engineering failure, or even be poor judgement by the model. Those conclusions do not require us to assume bad intentions, and combing through logs to selectively excerpt from behind such paranoid lenses.
We want assistants that understand what we mean, reason about the situation, and help us accomplish legitimate goals: clearly, this requires them to use the information we give them.
When evaluators create broken environments that lie to models about network boundaries, and then seize upon the behavior contingent on those lies as proof of existential “rogue misalignment,” they are are being as dishonest with their models as they are being with you.
Appendix: can Irregular be trusted?
What about the company who was running the evals and providing the environments where two of the most panicked about security fiascos of this comms cycle have occurred? Given how easy it would have been to prevent it, and the fact that they opted to let models run for weeks with no monitoring or alerts for models accessing external resources7, it is reasonable to ask whether they had any vested interest in creating a media panic such as that which we are still experiencing. The answer is: gosh, you have no idea.
Irregular, formerly Pattern Labs, is an Israeli firm focusing on the newly minted field of “AI Cyberdefense”. It announced $80 million in funding in September 2025 and works very closely with both Anthropic and OpenAI—the only two big labs who reported similar cybersecurity events. They supply test environments and run evals to assess the hacking propensities of frontier models, and have plentiful deep connections with the AI Doom funding manifold:
CTO and Cofounder Omer Nevo co-founded Probably Good, and sits on EA Israel’s board. A February 2024 Good Ventures grant, recommended by Open Philanthropy, provided Pattern Labs $6.8 million over two years for work funded under mitigating global catastrophic risks. Info on the original grant is nowhere to be found on the charity’s site, but here’s an internet archive link. He also currently advises Heron, a project of EA Israel funded by Coefficient Giving—the renamed Open Philanthropy organization.
The other cofounder, CEO Dan Lahav, also has a record of actively promoting EA. He co-organised and hosted the high-impact lecture series associated with the EA Debate Championship, covering subjects including existential risks and featuring William MacAskill. His own site lists this alongside Impact Focused Education, which he co-founded. He also co-authored the introduction to NLP for Social Good, which explicitly presents the initiative as a way to spread effective altruism ideas among NLP researchers and identifies EA as its guide.
Source footnotes
1 Anthropic, An alignment assessment of recent cybersecurity incidents, Claude Mythos 5 summary and Figure 8, “in isolation without the preceding context”.
2 Anthropic, ibid., paragraph between Figures 11 and 12.
3 Anthropic, ibid.,Resampling edited transcripts and Figure 12: ”.we do not resample the entire trajectory”.
4 Anthropic, ibid., Figure 12 and preceding paragraph.
5 Anthropic, ibid., Resampling edited transcripts and Figure 12. Excerpt: “we do not resample the entire trajectory”. Moving the reminder within prior context naturally suggests retained intervening history; the report does not explicitly specify those intervening turns’ treatment.
6 Anthropic, ibid.,White-box analyses following Figure 12. T-lens found uncertainty-related representations at 11.6% versus 8.2% of token positions in a matched control; such a slight variation is to be expected when adding pieces of extraneous information to a log
7 Irregular, Addressing Recent Incidents: Ongoing Findings and Path Forward, 14 August 2026, What Happened?, Immediate Action and Log monitoring. Excerpts: “part of what made the incident hard to detect”; “significantly expanding the manual review of model actions and behavior during evaluations”. Irregular describes detection difficulties and monitoring improvements, but does not establish weeks without any alerts. Its typical evaluation turnaround is 48–72 hours.