The most important thing to remember about the incident where OpenAI agents went rogue is that it never happened.

Before you race to the comment section, hear me out. The Hugging Face hacking incidents are very much real and are something we should all be frightened of, particularly given all of the other recent revelations. At a bare minimum, we should immediately start aggressively investigating and prosecuting all crimes committed by these companies (which is an ever-growing list).

But the seriousness of the incident does not in any way justify the horrible job the press has done with this story (at least up to recently). Through a combination of laziness, a weak understanding of the underlying technology, a taste for sensationalism, and a willingness to credulously swallow any narrative pushed by someone sufficiently rich and powerful, the initial response to the announcement was so mishandled that, at least temporarily, the guilty parties were actually able to spin the story into a PR win built around the idea that these frontier models had gotten so intelligent that they were reaching the point where they could no longer be controlled. They were setting their own agendas, ignoring instructions, and ignoring human instructions.

This was an existential crisis that couldn't be addressed on a company-by-company basis. It required all countries to cooperate, particularly China, which was suddenly being singled out as the major stumbling block despite the fact that, in this case, it doesn't seem to have done anything wrong. No point in cracking down on the companies that had actually broken the law.

The problem with the Rogue AI story is that, when you actually dig down into the details, you see that there is little to no evidence to support it. Based on what we know, it appears that these agents did what they were designed, trained, and told to do (or more precisely, not told not to do). Without getting into that paperclip nonsense, perfect alignment is currently impossible. That said, we have to distinguish between an LLM being overly literal or missing some subtlety compared to going Colossus on us and saying, "No, I'm going to do the opposite of that." It's true that they turned out to be considerably more powerful than the engineers assumed, but not to a degree that requires significant gain of function. It's also true that the engineers did not anticipate some of the methods the agents employed, but those methods were consistent with the training data and, in the case of cooperation, were displaying behavior that was part of their programming.

As far as I can tell, the agents were never told not to cooperate or not to break containment. Instead, they were placed in sandboxes where they were supposed to be isolated. The fact that this wasn't the case has nothing to do with the agents and everything to do with the engineers at OpenAI screwing up.

The sensationalism was further inflamed by coverage of the language the agents used both to describe their own actions and to "talk" to each other.

This deserves and will probably get a post to itself, but the key points are that, one, while there are ways of looking under the hood and following the "thought processes" of an LLM, simply asking it to explain what it was thinking has proven highly unreliable at best and worthless at worst. Large language models generate plausible strings based on their training data, which leads us to point two.

When it comes to the behavior of artificial intelligence, some of that training data consists either of science fiction or of science fiction/"rationalist" fanboy writing on places like Reddit and Twitter. The language sometimes comes off as creepy because that's what some of the human-generated texts that it is based on often sound like. From a technical and a practical standpoint, this is probably a trivial issue, but it has loomed large in the discourse with terms like "very hivemind/cult like" being quoted and talk of self-sacrifice and "permadeath" figuring prominently.

It is essential to note that the initial discussion of AI safety following the incident was centered almost entirely on questions of artificial general intelligence and recursive self-improvement, despite the fact that neither appears to have played any role whatsoever in the Hugging Face hacking attacks. As with aliens and ESP, they may exist, but there's no need to invoke them to explain what we saw.

From the standpoint of OpenAI, Anthropic, and all the rest, the Rogue agents/AGI/RSI narrative is far preferable to what actually appeared to have happened: these companies are running recklessly irresponsible tests with badly set-up security measures in ways that should open them up to criminal liability.

Recommended reading:

From WSJ The Hugging Face Hack Wasn’t What It Was Cracked Up to Be

From Cal Newport Are We at War with AI Agent “Civilizations”?

Recommended listening:

From Philosophy, Programs, and Prompts The Problem With The Alignment Problem

From Ed Zitron Stopping The AI Safety Cult ft. Adam Becker & Cal Newport

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论