Quick Thoughts on UKAISI’s “unsanctioned agent behaviour during cyber testing”
See the full UKAISI report here.
Some quick thoughts:
- Loss of control can happen during third-party auditing evals. The AI safety community usually thinks that loss-of-control situations arise from internal deployment within frontier labs, but this incident shows that they could also occur during third-party audits. Even though the likelihood may be lower––models are deployed within labs before being passed to third parties for auditing––smaller teams and relatively limited monitoring resources can still pose non-trivial risks. Note that this is not what happened at UKAISI as there was no sandbox escape. My point is more about the risk surface created by third-party audits, which to my knowledge has not been widely discussed.
- To some extent, I am not surprised that Mythos used social engineering to get humans to approve its attempt to add malware. First, 98% of cyberattacks reportedly incorporate some form of social engineering, and I believe models know this. It is therefore unsurprising that an attempted cyberattack included social engineering. In fact, social engineering worries me most because humans remain vulnerable to it (which is why phishing and scam calls still work), and the model needs only one human slip-up to succeed. As future security and monitoring stacks improve, I believe humans will increasingly become the point of failure.
- The ecosystem is not yet hardened enough for Mythos-level models, and I foresee such occurrences with abliterated open models. Three independent organizations (OpenAI, Anthropic, and UKAISI) found similar incidents, suggesting that these misaligned behaviors can emerge when guardrails are removed and models are subjected to cybersecurity evals. Although the incidents are not all in the same category (e.g., containment breaches vs social engineering), I am thinking about both sandbox security (i.e., whether the sandboxes have vulnerabilities) and eval protocols (i.e., how to ensure humans do not fall for more sophisticated social engineering attacks). While this suggests that currently deployed models with safeguards minimize the risks, abliterated open-weight models can still present a problem. In fact, Noah Lebovic has noted that Qwen3.6 would already bypass network controls during training to target real systems.
- I am worried about how future studies can realistically estimate worst-case uplift. People have argued that UKAISI’s cybersecurity evals should not be connected to the internet, but that would reduce their realism. My understanding is that UKAISI’s evals aim to mimic real-world deployments in which agents have internet access. More importantly, imagine similar misalignment arising during biological testing, where models without guardrails are evaluated in agentic virology settings, creating a chance that misaligned models could present severe biorisks. I think this possibility is much lower because biological attacks involve more friction (e.g., procurement and wet-lab synthesis) than cyberattacks, where the gap between intent and malicious action on the internet can be near-zero and the environment can change almost instantly.
评论
?
参与讨论