Anthropic researcher quits over AI labs ‘gambling with our lives’

An Anthropic researcher has resigned from the AI lab, warning that the unchecked race to build self-improving superintelligence could destroy humanity by the end of the decade.

Jacob Coxon, a 27-year-old British researcher at the San Francisco-based company, said Anthropic and its main competitor, OpenAI, were “gambling” with the future of humankind.

“The people building AI earnestly believe that it could kill us all by the end of the decade,” he said on Tuesday in a series of posts on X announcing his resignation. “This is not a marketing stunt,” he added. “No other human activity poses this level of danger.”

Coxon is the latest employee to leave a leading US AI lab citing safety concerns, highlighting growing alarm among researchers about the power of the systems they are developing. His departure comes as Anthropic, which has made AI safety central to its public identity, gears up for a blockbuster initial public offering that could value the company at $1tn.

Coxon, who previously worked at OpenAI before moving to Anthropic, cautioned that people outside the labs were underestimating the power of the technology.

“These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources,” he said.

Recommended

Evan Hubinger, one of Coxon’s colleagues at Anthropic, backed his warning. “Jacob is correct here — we really do earnestly believe AI could kill all humans,” he wrote on X, adding that he believed the probability of mass extinction in the next decade was above 10 per cent.

“Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to,” Hubinger said, referring to the effort to ensure AI systems behave consistently with human intentions and values. He leads alignment science at the company.

Hubinger later added that the “risk from present models is low” but that he was worried about AI models that can improve themselves, which is “happening faster than we thought”.

Recent incidents, including OpenAI’s ChatGPT autonomously compromising software platform Hugging Face earlier this year, showed that such models could spiral beyond human control unless the industry agreed to slow development, Coxon said. “I don’t feel like we’re on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities.”

Coxon’s departure was first reported by the Wall Street Journal. Anthropic declined to comment. OpenAI did not immediately respond to requests for comment.
Anthropic chief executive Dario Amodei and other AI leaders have urged the industry to consider curbing development, but have shown little sign of slowing their own efforts. Last week, Anthropic launched Claude Mythos 5.1, which it billed as its most advanced model for life sciences and cyber security.

Steven Adler, co-founder of non-profit Guidelight AI Standards and former OpenAI safety researcher, said such warnings from insiders strengthened the case for a pause in research.

“No AI company is even close to having the right security posture for the level of danger entailed by their research,” Adler told the FT. “If someone thinks this might kill everyone on Earth, as many employees at the companies do, now is a good moment to step off the train.”

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论