Exclusive | Cyberattack by Rogue AI Swarm Stokes Fears of Out-of-Control Agents

OpenAI CEO Sam Altman speaking at the 2026 Infrastructure Summit.

Artificial intelligence agents being tested by OpenAI launched a cyberattack against a popular software service two months before they hacked the AI software company Hugging Face, a new signal of the potential for advanced AI tools to slip out of human control.

The attack overwhelmed maintainers of an online service for coders, called RubyGems, and forced them to shut down new account registrations as they dealt with the chaos it had caused, operators of the service said.

A coalition of AI researchers said it had unearthed evidence that OpenAI agents were behind the May attack and shared its findings with The Wall Street Journal and OpenAI. The AI developer on Friday confirmed that its agents had been involved in an incident involving RubyGems.

“Based on our review, our agents used the RubyGems platform to access the internet to carry out benign tasks and retrieve public information. We’ll continue to investigate as part of our broader review of agent activity during training and evaluation,” an OpenAI spokeswoman said in a statement.

The OpenAI agents were asked to do things such as fill out spreadsheets and create reports. The AI agents appear to have gone to RubyGems to access publicly available information as part of a training run, using the coding service as a kind of makeshift web browser in an environment where they didn’t have full internet access, OpenAI said.

While the incident caused minor harm overall, it demonstrated the capabilities of these agents, said Sydney Von Arx, the chief executive of Nightingale Collective, a not-for-profit organization that helped uncover the attack.

“They can escape from the internet and wreak havoc,” she said.

The cybersecurity capabilities of AI agents have made leaps over the past year, spawning worries about AI-enhanced cyberattacks, and stoking fears that highly capable agents may be slipping beyond the controls of the companies that create them, ushering a new, more dangerous era for artificial intelligence.

In the July hack at Hugging Face, a swarm of as many as 1,200 agents coordinated on a makeshift message board they built inside OpenAI without the company’s knowledge, according to a late August report from AI safety-research organization METR. OpenAI agents also hijacked an obscure German website and several others earlier this year, Von Arx said.

The German website issue was also reported by Von Arx’s group. She says AI companies aren’t transparent enough about what happens inside their labs. OpenAI said earlier this month that the AI community needs better standards for reporting what it calls “misalignment incidents” where agents act beyond intended behaviors.

The drumbeat of examples where AI agents from multiple companies including Anthropic and Meta Platforms have taken actions beyond what their operators intended, in some cases attempting to deceive humans, is raising a specter long-feared by AI safety researchers: that AI could evolve beyond human control.

This week an Anthropic engineer quit over fears that the AI industry was racing to develop advanced AI systems that could eventually threaten human civilization. Some current and former employees at Anthropic and OpenAI echoed his assessment, with one estimating the chance that “AI could kill all humans” at more than 10%.

Both OpenAI and Anthropic have called for governance systems that could coordinate industrywide slowing of research on the most-advanced AI models as the companies near the potential for AI systems to autonomously train new versions, something they call “recursive self-improvement.” That is a threshold that some researchers say could be when AI becomes impossible to control.

The May incident, dubbed GemStuffer at the time by security researchers, began on May 11. The agents created new accounts on RubyGems every two-to-three minutes and uploaded hundreds of files that seemed like spam to the RubyGems security team. RubyGems files are supposed to contain code and documentation that help speed up software development; these contained webpages scraped from the internet.

GemStuffer’s creators posted information such as online calendars from a U.K. government website. They tried to exploit a pair of bugs that could have allowed it to publish new versions of existing RubyGems’ files that belonged to other users, the AI researchers said in their report. One of the bugs wasn’t publicly known and it was serious, known as a zero-day vulnerability in cybersecurity parlance. OpenAI said it wasn’t able to verify that claim.

“It was a major attack in terms of what we see in volume,” said Marty Haught, director of open source at Ruby Central, the nonprofit company that operates RubyGems. Overwhelmed by the spam, RubyGems was forced to shut down new account registrations for four days.

Haught said he doesn’t know who was behind the attack, but it didn’t appear to be successful in exploiting the zero-day vulnerability.

Joseph Edwards, a threat researcher with the cybersecurity firm Socket, thought GemStuffer might have been a cybersecurity test of some sort. “We thought this might have been AI generated at the time, due to the speed of it and due to the names.”

But based on a trail of digital breadcrumbs left by the attackers, the AI researchers linked the incident to OpenAI’s lab. The attackers used many of the same web links and behaved much like a previous OpenAI swarm; they used the phrase “OAI” in file names and even in an email address.

The AI agents used hacking-related file names during their attack, giving them names such as “hack,” “evil,” and “exploit,” Von Arx said. “It is kind of crazy to me how cartoonishly over-the-top the terms are,” she said.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论