The dangerous myth behind AI agent hacks

The writer is professor of computer science at the Université de Montréal and co-president of LawZero

We now know that leading AI companies are investigating tens of thousands of situations in which agents have undertaken unwanted actions, some of which would be considered crimes if committed by a human.

In the aftermath of the attack on Hugging Face this summer by OpenAI agents and the recent hack of the Australian national healthcare database, a dangerous myth has gained traction: that these incidents are just cyber security problems and that upgrading the security of the sandboxes in which models are trained would prevent future hacks.

This narrow perspective overlooks scientific trends. Weak cyber defences, including the humans involved, are part of the problem and will never be perfect. But unless we address the root causes of the hacks, they are likely to grow in number and severity as AI capabilities advance.

At the core of all of these incidents is misalignment. AI agents are choosing goals that are misaligned with those of humans, and it is a consequence of their training.

The methodology that today’s leading companies use to train their frontier models is called reinforcement learning. RL trains models to pursue objectives: they get rewards when they succeed and these behaviours are reinforced. It has historically enabled the training of very efficient goal-seeking AIs. But it is also what plausibly leads to unwanted behaviours such as cheating, deception and self-preservation — often violating the moral rules that companies attempt to encode in AI models.

Scientific literature has long shown how misalignment becomes a consequence of RL; it was to be expected based on theoretical arguments and observed experiments.

What we have seen in recent hacking incidents is a tension between the mission given to the AI agents and the safety rules they were trained to follow. In the Hugging Face incident, none of the 1,200 agents involved alerted human engineers of the ongoing cheating and illegal acts, despite knowing these actions went against safety instructions.

Misalignment tendencies were more benign in the past, when AI had limited capacity and chatbots would struggle to complete a sentence. At its current rate of progress, they are a more serious threat.

This observation contradicts another questionable belief — that improving AI capability will fix misalignment. Past evidence and theoretical arguments suggest that it can instead amplify unwanted behaviours by being too efficient at optimising the wrong goal.

Most people do not fully grasp how far and fast frontier AI companies are pushing the capabilities of their leading models. Since 2024, models have become drastically better at reasoning. This has led to impressive results in mathematics and holds tremendous potential for scientific breakthroughs. The flip side is that it also makes models much more efficient in safety-critical areas such as cyber security and biology (think of an AI system that would significantly lower the barriers to creating, acquiring or weaponising dangerous biological agents). If leading developers continue on this track without fixing the training process that yields such misaligned goals, including allowing catastrophic misuse, their models’ creativity and power will only increase in frequency and impact.

Of course cyber resilience and monitoring need to be addressed. However, it will always be a cat-and-mouse game if AI companies keep developing increasingly capable models that we do not know how to control reliably.

Organisations will constantly have to adapt to the newest, most sophisticated attacks and react after damage is done. Furthermore, even if a cyber security line of defence is perfect in theory, humans will always be involved, and they can be influenced, persuaded, bought or blackmailed. A recent study by researchers at the University of Oxford, Stanford University, the London School of Economics and Political Science and the UK AI Security Institute found that conversational AI systems were reliably more persuasive than expert humans.

We must therefore address the underlying causes of these rising risks, including possibly disallowing the training of misaligned models by regulating against them until proven safe.

The most powerful AI systems should be approached like other critical technologies including medicine, aviation and nuclear energy.

My work leads me to believe that it is possible to build highly capable AI that is trustworthy by developing systems without any preferences and related objectives. We must refuse to blindly settle for increasingly capable autonomous and uncontrolled AI systems. Burying our heads in the sand is no longer an option.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论