Anthropic CEO Says It’s Time to Slow AI Model Advances

Dario Amodei
Dario Amodei

Anthropic PBC Chief Executive Officer Dario Amodei said that the artificial intelligence industry needs to slow the pace of development for new models, citing growing concerns about “serious” risks to humans.

“We must slow the pace at which we improve the capabilities of AI models,” he wrote Saturday in a blog post. “Progress will still seem fast, and we must make wise use of the time we gain.”

He cited two main factors: AI’s ability to improve itself and the recent incident involving OpenAI and Hugging Face, where a swarm of AI agents collaborated to breach a third-party website.

“To be clear, pacing does not mean halting model training or technical progress, but ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this,” he added.

The anxiety about severe AI risks has begun to enter the mainstream, driven in part by this week’s high-profile resignation of an Anthropic researcher over existential fears that his company was acting irresponsibly.

Read More: Bridgewater’s Jensen Says AI Will Kill People Before It’s Curbed

Throughout the AI boom, executives at the leading AI labs have warned about potential existential risks – claims that have sometimes been dismissed as attempts to market the capabilities of their products and position themselves as the best stewards for the technology. Amodei himself has previously said there’s a 25% chance things go “really, really badly.”

Anthropic has long positioned itself as being focused on developing AI more responsibly and safely to mitigate the technology’s potential dangers and maximize the societal benefits. The company took pains earlier this year to limit the release of its Mythos model after determining it posed unique cybersecurity threats. Anthropic has also previously said the world needs a system to collectively decide when to slow work on the technology.

At the same time, Anthropic has remained locked in a fierce competition with longtime rival OpenAI to build increasingly sophisticated models that can automate more complex and valuable tasks for business customers. The two firms have both filed confidential paperwork to go public, with Anthropic expected to make its Wall Street debut as soon as this year.

Shortly after the Hugging Face security incident was revealed, Anthropic disclosed that its models had breached three organizations during cybersecurity tests, adding to concerns about the ability of AI labs to prevent their technology from running amok. Earlier this week, Anthropic said it had discovered a fourth hack as well.

While none of the agentic AI breaches, including Hugging Face, thus far have caused significant damage, Amodei said his worry is that “in 6–12 months such a swarm could be capable of taking over the entire internet with a persistent botnet (potentially causing hundreds of billions of dollars in damage), and that the scale of damage would continue to increase from there.”

Amodei wrote that Anthropic will commit to giving full access to third-party evaluators to verify safety practices and report incidents. He said Anthropic will soon bring these “embedded evaluators” into its offices and give them desks, badges, company laptops and “permissions mostly comparable to what internal risk assessment teams have.”

He said additional steps will require industry-wide coordination, as well as global cooperation.

“I continue to believe that AI can enormously improve the quality of human life. My desire to achieve these benefits is undimmed,” he wrote. “But the benefits will only be achieved if we build the technology in the right way, and — so long as we use the time we gain well — it is worth taking unusually deliberate care to get it right.”

Shortly after Amodei’s post, Elon Musk, who runs xAI Corp., wrote, “Dario is right.”

Growing Alarm

Earlier this week, OpenAI CEO Sam Altman said his company is considering slowing down the development of cutting-edge AI, preferably in conjunction with the broader industry.

The company’s top scientist, Jakub Pachocki, had posted a warning about the dangers of AI, saying that he believes companies should be “coordinating to slow down future development as needed.”

On Tuesday, Jacob Coxon, an AI researcher, quit his job and accused both of his former employers — Anthropic and OpenAI — of “gambling with our lives” by racing toward superintelligent AI. Coxon said in a social media post that the people building AI believe it could “kill us all by the end of the decade.”

In response, Evan Hubinger, a current Anthropic employee, said he and others at the company do worry about this scenario. “I personally think it is >10% within the next decade,” he wrote on X.

Back in July, more than 1,000 staffers across top AI firms signed a petition calling on the US government to support a mechanism that would help “deliberately pace” AI development to prevent the technology from advancing too fast.

AI models are proving increasingly adept at carrying out their own hacks, with Anthropic, OpenAI and Meta Platforms Inc. all disclosing in recent months that their models had escaped testing environments, accessed the open internet and breached real-world victims during testing.

Amodei, in his post, wrote that a coordinated strategy would allow AI industry leaders in the US to complete needed safety work without sacrificing any competitive advantage. He said that this would require some antitrust exemptions in the US to allow companies to coordinate in specific areas, as well as cooperation with China.

“If we greatly restrain our AI capabilities in the belief that China will do the same, and then China defects, AI could be so powerful that such a defection could lead to their geopolitical dominance,” he wrote.

So far, however, the Trump administration has shown little interest in flexing regulatory muscles to put guardrails around AI development.

Read More: Trump Brushes Off AI Doomsaying to Safeguard US Lead Over China

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论