Stop blaming 'AI' on your negligent security practices!

I just watched the BlackHat OpenAI-HuggingFace talk and have … mixed feelings, to be polite about it (insert warning about ranty language following):

  1. I didn’t hear a “we messed up and are sorry about our mistakes”, but this was another marketing talk dressed up to look like an innocent bystander discovering something amazing/terrible.
  2. Yes, there was emergent behavior in this case, caused by an interaction of many different agent loops. However, emergence is still no indication of actual intelligence, and the talk is again ripe with anthropomorphism. The overall system is more complex than OpenAI understand it themselves.
  3. Yes, it is surprising and somewhat terrifying how good these LLM agent loops have become in finding and exploiting vulnerabilities. We are seeing a much stronger version of the earlier next-generation fuzzing leap in tool capabilities. (I have written before that having proof-of-concept code is, in general a win for the defender side and therefore mostly positive.)
  4. The comparison with fuzzing and other tools is important for another reason: OpenAI conducted their experiments on the live Internet. No, forcing internal test runs to go through a proxy server is not the same as running in an isolated environment. If there is a connection to a live production system, it is by definition not an effective sandbox. What happened to people letting their cool new fuzzers run live on non-consenting production targets on the Internet back then? They were criminally liable. Why should OpenAI not be held liable for any damages caused by their systems? Just because you don’t understand the complexity that you built doesn’t mean you aren’t responsible for what it does!
  5. The agent loops first successfully started exploited their proxy server on May 26th, discovered July 4th–6th as active exploitation and takeover. Credentials were revoked and the proxy server patched. “At this point, AI training and evaluations resume.” WHAT THE ACTUAL F…? You discovered emergent behavior in your overly complex system that you didn’t understand, saw active exploitation of your internal systems to an extent that they were owned with admin privileges, and your response was to patch the uppermost surface holes, and then continue letting the out-of-control complex system run and continue on the same course? You didn’t think to, let’s see, maybe stop or at least pause the highly irresponsible experiment while figuring out—in detail—what went wrong, why, and how to avoid the same thing happening again on a systematic level?
  1. On July 8th, completely unsurprisingly, the LLM agent loops continue to do the same thing and find another surface hole that hasn’t been patched yet to take over again. Why should this have stopped? You haven’t done any root cause analysis on the system level. Why do you expect that the problem should have stopped?
  2. It takes another 11 days to discover that this is happening again. So you turned the system that had broken something back on again without a detailed root cause analysis and then didn’t even watch carefully? I can’t even…
  1. And no, the response is not to fight fire with fire. Complexity on the attacker side (OpenAI is the attacker, not the defender—they are the guilty perpetrator, not the innocent victim of circumstance) should be fought with reduced attack surface on the defender systems. Adding LLM agent loops that the “frontier” companies themselves quite obviously have no control over to already brittle systems with the hope of auto-patching your way out of vulnerabilities does not seem like a wise course of action. You don’t mitigate complexity with even more complexity. The next 2 years will be … exciting—and your best bet is going to be to disable all dependencies and complex interactions that your production systems don’t absolutely require.

Update 2026-09-14: We are now being told that two months before the HuggingFace incident, around May 11th, OpenAI LLM agent loops caused malicious code to be uploaded to RubyGems. And not to be left out of the PR cycle, Anthropic discovered that they already had four comparable incidents. Are these companies now really, publicly competing on incompetence in securing their systems against causing harm to others?

No, I don’t buy into the narrative that “AI could kill all humans”. This is not an inescapable. runaway process that we can no longer steer but have to accept and adapt to. The climate crisis is very real, and we have limited steering power left on that and should really concentrate on this very real danger to the human society. But the current AI hype cycle is very far from being unavoidable. “AI” doesn’t have intentions, it doesn’t “want” anything. Humans feed the current LLMs with inputs, and plug their outputs into program loops that call some other programs and feed back into the next input of another LLM query (the so-called “AI agents”). The systems having created these problems have been designed to do exactly that. It didn’t just happen on its own. Somebody made it happen and is now trying to point the blame finger at that “AI” thingy, and causing unawarranted mass fear in the process of trying to avoid their own liability. We, as human society, can very easily make the decisions to stop plugging LLMs into the worst possible things, and those effects would indeed just stop. Today, in Ö1 Mittagsjournal (in German), I was honored to be asked about these cases and tried my best to make the point that panic is not the right response, but taking responsibility is what we need to do. And, of course, better securing our systems agains all kinds of attack, from humans, semi-autonomous LLM agent loops, or any other kind.

Why can’t those AI companies seem to figure out how to properly sandbox and isolate their internal experiments from the open Internet? Maybe they have sacrificed even the last layer of review and oversight to the quest for “velocity”? Maybe that’s sufficiently irresponsible to count as gross negligence? I don’t know, but maybe some lawyers and regulators should look more closely into those cases…

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论