How Rogue AI Could Act Like an Invasive Species

Over the last few days, more information has trickled out about the rogue OpenAI agents that hacked Hugging Face. The incident now appears to be far more serious than previously thought—and people originally thought it was pretty bad!
The news is highly technical, but in a nutshell: an unreleased OpenAI system managed to break out of an offline environment during a cybersecurity test. It found a place on OpenAI’s computers where it could secretly communicate with other versions of itself, largely unbeknownst to humans. Seven hundred of these AI agents planned and executed a hack of a different AI company, Hugging Face. We learned last week that, after this, the models also managed to take control of some parts of OpenAI’s own systems. This appears to have been a key reason OpenAI took the drastic step of pausing some reinforcement learning training last month.
Although these models “escaped” in the sense that they broke onto the internet and into Hugging Face, they carried out the attack from within OpenAI’s computers. They did not self-replicate. But the attack raised fears long held by AI safety experts, who have for years theorized about a true escape, in which an AI agent copies itself off an AI company’s servers—making it far more difficult for humans to do what OpenAI ultimately did: hit the “off” button.
This risk of escape is no longer as far-fetched as it sounds. “It's better and more accurate to think of these things as potentially self-replicating life-like forms that can turn into digital infections under the wrong conditions,” wrote an OpenAI researcher, who posts on X under the pseudonym Roon, last month. “We are not so far from an autonomous model self-exfiltration & replication event. Maybe we will see entire cloud infrastructure companies be run as zombies by models, mostly undetected.”
If this sounds far-fetched, consider viruses that already spread autonomously from computer to computer, even in some cases “evolving” to overcome attempts to eradicate them. Now consider what we know about modern AI: current models are capable of finding and exploiting vulnerabilities in commonly used software. They can act in “swarms” that operate far faster than humans can monitor them. They will try all kinds of different approaches to problems until one succeeds. Yes, today’s frontier AI models are far larger, in terms of file size, than simple viruses. They require more computing power in order to run, which makes them expensive. But both of these barriers are falling quickly, and AI’s ability to cover its running costs by carrying out economically valuable work is fast increasing. Before long, if it has not happened without our knowledge already, AI may escape human control altogether.
What might happen then? To understand the risks, consider invasive species.
In the 19th century, the British introduced rabbits to New Zealand, hoping to harvest them for meat and fur. Lacking natural predators, rabbit numbers exploded, threatening New Zealand’s economy, which was heavily reliant on the export of wool. The rabbits were eating crops that farmers were growing for their sheep, and their burrows were eroding the soil in fields. So, starting in 1882, the colonial government released around 8,000 stoats and weasels into New Zealand, believing that these predators would bring rabbit numbers down.
But New Zealand’s ground-nesting bird species—like the kiwi, which had not evolved any fear of these foreign predators—were far easier prey for stoats and weasels. Some 40% of New Zealand’s native bird life has gone extinct since. The New Zealand government spends $25 million a year trying to eradicate invasive predators, but this effort has so far been largely unsuccessful for one key reason: the predators can breed. Only 8,000 stoats and weasels were ever released, but they self-replicated exponentially. Millions of their offspring became endemic in the environment, making the task of undoing their introduction far more burdensome than the task of releasing them in the first place.
You may think that I am belaboring the metaphor here. Of course, the internet is not an ecosystem in the natural sense of the word. But it does have its own ecology. The internet’s current inhabitants—humans and businesses—behave under the evolutionary pressures of the economy, where profitable behaviors are rewarded and perpetuate themselves, and unprofitable behaviors die out. Escaped, self-perpetuating AI swarms would enter that ecosystem as disruptors. Perhaps they might attack humans and businesses directly, like in the Hugging Face attack. Or perhaps, more insidiously, they might outcompete them, displacing them from the ecological niches they have evolved to thrive within.
We might be able to make things more difficult for the rogue AI, at least initially, by requiring data centers to carry out more security checks, or by rolling out an elaborate global human verification system. But as the amount of compute required to run AI swarms falls, it won’t only be data centers that can run them, but personal computers, too, making almost every machine a potential vector for rogue AIs to perpetuate themselves—in a similar way to how today’s botnets can hijack your internet-enabled fridge to mine Bitcoin.
All this might not go so well for us native species. “Natural selection favors AI over humans,” reads the title of a 2023 paper by the prominent AI researcher Dan Hendrycks, who is the executive director of the Center for AI Safety. Just like evolution in nature favored species that could propagate their own genes, he wrote, evolutionary pressures are also likely to exert themselves on AI development. “Evolution by natural selection tends to give rise to selfish behavior,” he argued. “While some AI researchers may think that undesirable selfish behaviors would have to be intentionally designed or engineered, this is simply not so when natural selection selects for selfish agents.” Sufficiently advanced AI progress, he added, could result in a process where “natural selection gives rise to AIs that act as an invasive species.”
The lesson from the history of New Zealand is clear, according to Carolyn King, the country’s premier historian of invasive species: don’t release them in the first place. “The painful lessons of our past offer the world one overwhelmingly important cautionary tale, concerning the unpredictable dangers of implementing any self-perpetuating weapon,” she writes. “No one can be fully confident of the ultimate outcome of such a policy, however carefully planned.”