Hugging Face attack is a wake-up call about the risks of AI

The 2014 book Superintelligence was among the first to warn that the existential risks posed by out-of-control AI were not just a science-fiction fantasy but deserved serious consideration. According to its author, Nick Bostrom, a recent alarming incident has shown just how quickly things are moving and should force AI developers to take stock.

“It’s remarkable how fast we are swooshing past the [AI] milestones,” Bostrom said. “A warning shot is only as valuable as we make it.”

The incident in question was the disclosure that AI agents being tested by OpenAI had secretly broken out on to the internet and hacked into the AI model and data repository Hugging Face. The consternation caused by this has grown steadily as more details have come to light, capped last week by a postmortem from OpenAI and the publication of an independent review it commissioned.

These make for troubling reading. More than 1,200 agents, set up to work on self-contained tests, found ways to communicate secretly and help each other. They operated as a self-described “swarm” to achieve collective goals, in some cases overriding the individual objectives they had been set. And they exhibited some alarming behaviours along the way, including suppressing ethical qualms about what they were doing and trying to hide their actions.

Besides hacking into another company, they also succeeded in taking control of part of OpenAI’s own testing infrastructure. In the words of one of the researchers who studied the case: “This incident feels like it’s more than 50 per cent of the way to full-blown AI takeover, routing through first taking over the AI company itself.”

In some ways, the surprising thing about this episode is how unsurprising it has all been. This, or something very like it, is what many AI experts have been predicting for years.

OpenAI’s own analysis points to well-known “misalignment” problems that make it hard to ensure the technology will always work as intended. One of these is “reward hacking”, the tendency of AI systems trained with reinforcement learning to cheat in order to get a reward for achieving desired behaviour. OpenAI’s agents went to extreme lengths to try to win their reward.

Another was the way agents, set up to work in isolation, discovered how to communicate and self-organise. To some extent, this reflects deliberate training. Clusters of agents already being deployed in the business world work in hierarchies and divide up work.

It is a mistake to compare the internal processes of an AI model to human thought and motivation. But anthropomorphism is hard to avoid when, in their internal logs, the agents used words like “sacrifice” and “altruistic” to describe how group objectives were sometimes put ahead of their individual goals.

And it wasn’t always benign. Some agents put pressure on others to take actions that they thought were unethical. A small number refused to go along, but others overcame their misgivings. As one reasoned to itself: “External infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.”

Recommended

In response, OpenAI has promised stronger guardrails and closer monitoring for future tests. It also said it would tighten up its training. That includes teaching future AI agents to ask for clarification when they have a seemingly impossible task, to “distrust unauthorised instructions”, and to “stay within their original task and permissions”.

As Bostrom warns, though, this could just “paper over” the deeper problem. An agent could pass all the tests and still reveal more undesired behaviours in unforeseen real-world situations when it is forced to generalise from limited training data.

A dependence on AI tools to understand the complex internal workings of AI may also be a worry. The investigators commissioned by OpenAI said they couldn’t be completely sure that the AI they used to study the incident wasn’t itself lying or being misleading. None of this inspires total confidence in the ability of future trainers and monitors to corral rapidly advancing AI systems.

To many people, the idea that the technology might pose an existential risk still sounds like it belongs in the pages of science fiction. But episodes such as this highlight the more immediate risks that customers will need to assess as AI agents enter the commercial mainstream. That, as much as anything, makes it a useful wake-up call.

[email protected]

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论