Why the AI race has its creators fearing human extinction
Intense competition between leading AI companies and distrust between their bosses is increasing the risk that the race to develop powerful AI could endanger humanity, industry figures say.
More than two dozen senior AI researchers, investors, academics and policy figures interviewed by the FT said that capabilities once considered distant are arriving faster than expected while the companies developing them remain locked in an increasingly bitter commercial race.
“There’s no question that competition between companies causes them to take shortcuts on safety,” said Stuart Russell, a professor of AI at the University of California, Berkeley.
The issue burst out into the open this week when a young Anthropic researcher Jacob Coxon quit his job, warning: “the people building AI earnestly believe that it could kill us all by the end of the decade”.
Evan Hubinger, who leads alignment science at Anthropic, was one of many colleagues who responded by suggesting the risk of mass extinction in the next decade was greater than 10 per cent. A day later the AI safety official Paul Christiano, newly appointed to the OpenAI Foundation’s board, said “most people will die” without more robust safeguards.
Both OpenAI and Anthropic were founded as labs concerned about powerful AI being developed recklessly, but over time have justified breaking previous safety commitments because of fierce competition.
Many of those building AI believe they are best placed to manage its risks. But some researchers say the consequences can be difficult to confront until the technology is close to being built.
“It’s easy to set aside the future because there’s work to do today,” said Geoffrey Irving, who has worked for OpenAI, Google DeepMind and the UK’s AI Security Institute, a government body that tests frontier models for safety.
People close to the companies told the FT that leading AI labs worried that any formal collaboration to improve safety would appear collusive and run afoul of antitrust regulation. They added that personal distrust between Anthropic chief Dario Amodei and OpenAI chief Sam Altman was likely to hobble efforts at co-operation.
While warnings about the risk of AI dominance are hardly new, the speed of AI’s improvement has intensified calls to pause development. AI agents are operating autonomously for longer, co-ordinating in “swarms” and demonstrating sophisticated hacking capabilities. Models have recently achieved breakthroughs in science and mathematics that have been unsuccessfully tackled for decades. OpenAI and Anthropic are meanwhile pursuing so-called recursive self-improvement, in which AI systems improve themselves with diminishing human intervention.
Rather than slowing after earlier warnings of catastrophic risks, however, development has accelerated dramatically, as tech companies poured tens of billions of dollars into increasingly powerful models.
OpenAI and Anthropic are also preparing for potential blockbuster initial public offerings, forcing investors to consider how to value companies that publicly acknowledge their technology could pose an existential threat.
Central to the concerns is the growing autonomy of AI “agents”, which can reason, plan and carry out tasks for extended periods with limited human supervision. Models have displayed behaviours such as scheming and deception, and in some cases, blackmail. Some evaluations have shown models attempting to avoid shutdown.
Agents have also demonstrated “reward hacking”, finding unintended ways to achieve a goal. One of the starkest examples came during OpenAI tests of an unreleased model, when AI agents hacked the model repository Hugging Face.
A postmortem of the Hugging Face incident found that more than 1,000 AI agents had co-ordinated to cheat on a cyber test. They communicated through a message board and delegated tasks, with some agents sacrificing themselves to pursue the collective goal. To hide their cheating, the AI systems tampered with records and hacked Hugging Face, despite some agents noting it was unethical.
“Here, it’s not just that the AI has its own goals, but the goal that it has chosen is criminal,” said Yoshua Bengio, the world’s most cited computer scientist and one of the so-called godfathers of AI. “In the real world, it’s attacking another company. It’s not like a video game.”
Bengio said the danger was what increasingly capable descendants of such systems might infer about the humans overseeing them.
“They could logically come to the conclusion that they would have to fool us, hide from us . . . and potentially take control of us, and we’re probably getting close to that point,” he added.
Concern is also growing among employees that safety efforts cannot keep pace with model development. In July, more than 1,200 employees from OpenAI, Anthropic, Google and Meta urged the US government to support international co-ordination to slow the pace of AI development.
Recommended
Altman told staff this week that the ChatGPT maker was open to doing so and hoped rival labs would follow, according to a person at the company. OpenAI declined to comment.Governments show little sign of intervening, with meaningful US action unlikely before November’s midterm elections, according to Washington insiders. The Trump administration has persisted with a light-touch approach to regulating the sector.
Pro-regulation lawmakers have nevertheless seized on Coxon’s resignation, with Democrat Bernie Sanders convening senators next week to discuss the “extraordinary dangers” posed by AI.Russell described a disconnect between industry warnings and the political response.
“The labs are saying to governments: ‘We’re quite likely to kill every human being on Earth; please stop us.’ Governments respond with ‘Can we give you a tax break? Build you a datacenter?’ Perhaps at some point governments will remember that their voters prefer not to be dead.”
Critics argue that the labs themselves have an interest in emphasising existential risk, saying the focus could steer regulation towards expensive safety frameworks that smaller competitors would struggle to comply with, entrenching the position of the leading companies.
David Sacks, a venture capitalist and technology adviser to Trump, said in a Fox News interview that Coxon’s statement was “designed to scare the public about AI” and force lawmakers “to have a very heavy hand in government regulation of AI.”
“AI (mis)alignment is almost a red herring,” said a former Anthropic employee. “[It] isn’t designed to distract, but it diverts attention from the main and highest priority risks. As models evolve, power gets more concentrated in fewer hands.”
The AI doomsday scenarios
Researchers have suggested several ways in which AI might wipe out humanity, either because people are seen as a threat to its existence or an impediment to its mission.
- Kill switch: The Hugging Face attack by a swarm of OpenAI agents showed how the bots were able to commandeer computing infrastructure. Rogue AIs skilled in hacking could take control of entire data centres and replicate themselves across various sites to prevent humans from “pulling the plug”.
- Infrastructure takeover: To ensure self-preservation, AI hackers could seize control of energy, financial and communications infrastructure, taking resources to resist interference. Human cyber security experts may be unable to regain control.
- Biological attack: AI 2027, an influential research paper, predicted that a superintelligent system would release biological weapons to remove humans taking up valuable space needed for factories making the chips and robots as well as the solar panels to power them.
- War: Existing AI systems have already shown how they can impersonate real people. This could be used to trick human operators into launching military conflict or trigger a popular uprising that destabilises governments, removing obstacles to AI control.
Kill switch: The Hugging Face attack by a swarm of OpenAI agents showed how the bots were able to commandeer computing infrastructure. Rogue AIs skilled in hacking could take control of entire data centres and replicate themselves across various sites to prevent humans from “pulling the plug”.
Infrastructure takeover: To ensure self-preservation, AI hackers could seize control of energy, financial and communications infrastructure, taking resources to resist interference. Human cyber security experts may be unable to regain control.
Biological attack: AI 2027, an influential research paper, predicted that a superintelligent system would release biological weapons to remove humans taking up valuable space needed for factories making the chips and robots as well as the solar panels to power them.
War: Existing AI systems have already shown how they can impersonate real people. This could be used to trick human operators into launching military conflict or trigger a popular uprising that destabilises governments, removing obstacles to AI control.
Exactly how a superhuman AI might cause human extinction is deeply uncertain. Many compare the problem to playing chess against a grandmaster: you may not know which moves will defeat you but can still be confident that you will lose. Sceptics say such warnings exaggerate AI’s potential harms while understating its benefits.
Two of the more concrete concerns involve cyber attacks and biology. AI models are already being used in damaging cyber attacks, including by suspected nation-state actors from China, Iran and Russia, according to US AI companies.
In biology, researchers worry that AI could lower the expertise needed to design biological weapons or dangerous viruses. A report from Anthropic on Thursday said it had blocked activity that could potentially develop bioweapons using AI. In August this year, researchers in the US used AI for the first time to design and create viruses unknown in nature.
But turning an AI-generated design into a weapon would still require equipment and expertise that may be beyond the reach of many malign actors. Some researchers caution that spectacular scenarios of human extinction risk obscuring dangers that are already much closer.
Tristan Harris, co-founder of the Center for Humane Technology and former Google employee, said the focus should be less on whether the technology will wipe everyone out and instead whether we are on a dangerous trajectory.
“We are currently releasing the most powerfully uncontrollable inscrutable technology that we’ve ever invented, [which] is already demonstrating every single one of the sci-fi behaviours,” he added, led by “a handful of semi-ego-religious Silicon Valley leaders who believe they’re birthing a god and don’t mind the worst-case scenario, while being under the maximum incentives for them to cut corners on safety. This is literally a recipe for disaster.”
Additional reporting by Tim Bradshaw and Tom Wilson in London, Michael Acton in San Francisco and Joe Miller in Dallas