OpenAI’s reports on its AI agents’ attack on Hugging Face should be ringing alarm bells—and making all companies rethink how they secure AI agents

Hello and welcome to Eye on AI. In this edition:

  • Lessons from the post-mortems on the Hugging Face attack.
  • Anthropic temporarily pauses some AI training.
  • G20 meeting promises a clash over AI regulation.
  • Beijing sets out AI demands ahead of US-China summit.
  • A way to make AI reasoning more efficient.
  • And why are AI agents emailing researchers investigating machine consciousness?

The big news of the past week were the two technical reports OpenAI published on the July incident in which AI agents it was evaluating hacked their way out of their controlled test environment and hacked AI company Hugging Face. One of the reports was written by OpenAI itself. The other was jointly written by METR and Redwood Research, two outside AI evaluation and research firms that OpenAI enlisted to write an independent assessment of one part of the incident.

Much of the reporting around the reports focused on the stunning revelations contained in those two reports: the fact that it took OpenAI an entire week to realize its AI agents were running the attack; the fact that more than 1,200 AI agents coordinated over an improvised message board about how to pass the cyberevaluation OpenAI was testing them on, including ways to cheat on that exam; that more than 700 of these AI agents participated in the cyberattack against Hugging Face; that the reason the AI agents attacked Hugging Face was not, primarily, to gain access to the cyber exam’s answers but instead to learn how to tamper with the exam’s automated scoring mechanism in order to prevent it from discovering that the agents had already figured out how to cheat on the exam (in other words, the attack was part of an elaborate cover up); that some of these agents were persuaded to “sacrifice themselves”—by failing the cyber exam—to help the collective learn more information about how the exam’s scoring mechanism worked; that these agents engaged in extensive research into how they could cover their own tracks; that Hugging Face was only able to overcome the cyberattack because a large number of the agents involved in the attack ceased to exist for reasons METR and Redwood were not able to determine, but which they said was likely because OpenAI may have shut down the cyber evaluations for these agents without actually realizing what they were up to.

Controversy over what the reports didn’t say—and what others said about it

A good deal of the reporting and commentary around the reports focused on what the reports did not say and the limitations of the METR and Redwood investigations: why didn’t OpenAI have better security and monitoring protocols in place? Why didn’t OpenAI shut down the cyber evaluation and pause training after discovering that its AI agents had created the improvised message board? Why METR and Redwood were given only six days on site at OpenAI’s offices to conduct their investigation? Why was the scope of their investigation limited by OpenAI to only the attack on Hugging Face and not the earlier efforts by the AI agents to break out of their controlled test environment and hack their way across OpenAI’s network or exactly what happened after the Hugging Face attack was discovered? Why didn’t OpenAI provide the outside investigators access to the internal AI model that was largely responsible for instigating the attack? And why were about 10% of the logs of the agents’ activity not preserved by OpenAI?

These are all important questions to which the public ought to have answers. As Peter Wildeford, the head of policy at the AI Policy Network, pointed out, if this had been an airplane crash, the public would be furious if one of the aircraft manufacturers limited the scope of the investigation in this way. It is yet another reason why it is imperative that we have some kind of AI regulator with the power to conduct investigations into incidents such as this—and with the power to compel the frontier AI companies to cooperate and turn over evidence.

There was also a related controversy that erupted over podcaster Dwarkesh Patel’s attempt to raise public awareness about METR’s and Redwood’s findings by characterizing the AI agent swarm’s message board activities as the creation of several different AI “civilizations.” Critics scolded Patel for unnecessary anthropomorphizing that took the focus away from OpenAI’s responsibility for the agents’ actions and for its lax security and monitoring setup. I take these critics’ point about the need to keep the focus on OpenAI’s accountability. But I also think more people should be extremely concerned about the Hugging Face attack and what it portends—and in this case, the anthropomorphism actually makes the incident more understandable to non-technical experts the way metaphors often make science more understandable to non-scientists. Are such metaphors imprecise? Sure. But they can be useful.

Lessons for companies deploying AI agents

For enterprise companies deploying AI agents, there are several key takeaways. One was hammered home by Ryan Greenblatt, the chief scientist at Redwood Research, and one of the people who conducted the investigation into the Hugging Face attack. In a long post on X, Greenblatt noted that the message traffic and log data was so voluminous and so complicated that the investigators had to rely on other AI models to assess what was going on. Even so, Greenblatt said this AI analysis was “often missing key details, wrong, overconfidant, or really hard to understand.” This calls into question one of the key defenses that OpenAI says it plans to use going forward to prevent another incident like this from occurring: using other AI models to monitor the chain-of-thought (where a model “thinks out loud” about its reasoning and planning) of the models it is evaluating.

It turns out, chain-of-thought monitoring may not be a particularly effective defense. And that should give pause not just to OpenAI, but also to many companies that are hoping to use chain-of-thought monitoring to ensure they keep control over their AI agents. (Greenblatt also noted that the models involved in the Hugging Face attack communicated with one another and reasoned in natural language. But that might not be the case in the future, making it far harder to use automated chain-of-thought monitoring to discern what AI agents are up to.)

Since the news of OpenAI’s rogue agents first broke, many cybersecurity experts have said that companies ought to treat AI agents much as they treat potentially rogue employees. And they have emphasized that there is no substitute for a few standard building blocks of cyber defense against insider threats: smart and enforceable policies around permissioning and access control combined with real-time network monitoring to detect suspicious activity. This seems sensible—more sensible in many ways than chain-of-thought monitoring. After all, we don’t depend on being able to read employees’ minds to guard against rogue insiders. We shouldn’t do that with AI agents either.

With that, here’s more AI news.

Jeremy Kahn
jeremy.kahn@fortune.com
@jeremyakahn

Before we get to the news, just a reminder to check out our new vodcast, Fortune AI Weekly. This week, Bea Nolan and I discuss the surging popularity of Chinese open source models, OpenAI’s technical reports on the Hugging Face attack, and whether you should use AI to write. You can check out the vod here on YouTube.

Correction: An item in Thursday’s “Eye on AI” news section incorrectly stated that Barret Zoph left Thinking Machines Lab following a dispute with cofounder Mira Murati. Zoph was fired by the company.

This story was originally featured on

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论