Anthropic’s Moral Conflict Is Playing Out in Real Time
Dario Amodei built Anthropic to avoid this very kind of moment.
In a stunning essay Saturday, the chief executive of one of the leading AI companies echoed many of the doomsday warnings—from current and recent employees—that sparked panic across the nation this past week.
“Along with my co-founders and employees, I have grappled with this duality of risk and benefit since the beginning of Anthropic,” he wrote. “Not building the technology deprives humanity of benefits or simply places AI in the hands of authoritarian powers, while building it too fast is reckless.”
He reiterated a call for the AI industry to slow down development of advanced artificial intelligence to give it enough time to ensure it is being done safely. But, just weeks ahead of what might be the largest IPO ever, Amodei stopped short of doing the one thing in his power: pausing or slowing Anthropic’s own work.
It’s a classic prisoner’s dilemma.
If all AI labs, including bitter-rival OpenAI, agree to slow their work, it could well benefit all of humanity—if Amodei’s warnings are to be believed. That slowdown likely gives Anthropic, already a leader in the field, an advantage.
But if Anthropic is the only one to heed the call, it falls behind OpenAI.
Even if OpenAI goes along with Anthropic, as its CEO, Sam Altman, suggested on Saturday that it would, the U.S. could fall behind China.
Then there’s the money. The trillions of dollars at stake that will make people like Amodei among the richest in the world and pour fuel on Anthropic’s ability to further accelerate development of its AI.
In June, the company filed confidential paperwork ahead of an actual listing. The company was targeting going public this month or October. It is looking to raise as much as $100 billion at a valuation of about $2 trillion, ahead of a planned OpenAI initial public offering, The Wall Street Journal has reported.
Yet, given this past week, it is hard to believe Anthropic is really ready to become a publicly traded company in coming weeks.
Not when its own employees are saying publicly that they earnestly believe there’s a chance their technology will wipe out humanity in the next few years. And they haven’t figured out how to prevent it.
Understandably, those statements, coupled with recent rogue AI hacks, have sparked great concern. Dozens of lawmakers—Republicans and Democrats alike—urged action, including halting AI development.
But the really scary thing about the doomsday talk wasn’t the threat of Anthropic’s AI, much of which we’ve heard before. Rather it was the suggestion that the unique safeguards Amodei built to resist market pressures at Anthropic are failing.
“At Anthropic, the stakes are well-understood, but they are locked in a race to get there first—they believe no one else will act responsibly, so they must do it themselves, despite the risk,” Jacob Coxon warned on X Tuesday.
The AI researcher, who had previously worked at OpenAI, first announced he was leaving Anthropic in an interview with the Journal and subsequently appeared on CNN and Fox News. His viral story took AI doomerism—long a mainstay of Silicon Valley debate—mainstream in a way that surprised the industry.
What made his warnings so much more powerful was that the sentiments were endorsed by his former colleagues, including Evan Hubinger, a current team lead at Anthropic who gave a greater than 10% chance of AI killing all humans within the next decade.
“I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to,” Hubinger posted on X in response.
For a moment, let’s put aside conspiracy theories about the motives of Anthropic’s AI warnings, and take them at face value.
Then, we’re watching, in real time, a workforce wrestling with the moral conflict of what it means to no longer be a science experiment and instead be in the business of making AI that, many of them believe, can be dangerous.
Around the same time as the IPO filing, Anthropic published a lengthy essay arguing why a global effort to pause AI development was a good idea. “Without a global coordination mechanism, companies and governments will have to make difficult decisions about safety while under competitive and geopolitical pressures,” the company said.
Weeks later Amodei and others across the industry signed a letter calling on the U.S. government to support an international effort to pace advanced AI development.
A pause hasn’t happened. And the race continues.
The money at stake is why some in tech don’t want to take Anthropic’s warnings seriously. They chalk it up to marketing wrapped in false alarmism. Or suggest it is part of another foreign influence campaign to put the U.S. behind China. Or say it is another piece of a sophisticated lobbying effort to ensure regulations are written to cement Anthropic’s market-leadership.
Much of their frustrations over the past week’s flashpoint stem from the fact that for years such doomerism has been a part of Silicon Valley’s debate around the technology.
And there’s probably no bigger p(doom)’ers than the folks at Anthropic. It was founded in 2021 by Amodei and a handful of others who left OpenAI out of concerns that the AI lab wasn’t taking safety seriously enough in its efforts to quickly develop the technology.
Amodei has warned that AI could lead to Great Depression-like job losses and, just a year ago at an Axios conference, said there was a 25% chance that the future of AI will go “really, really bad.”
Such advocacy is in keeping with the company’s beliefs. In its own description, Anthropic is an AI safety and research company set up with a unique governance structure aimed at protecting it from the whims of investors.
The setup, Anthropic has said, “can ensure that the organizational leadership is incentivized to carefully evaluate future models for catastrophic risks … rather than prioritizing being the first to market above all other objectives.”
Yet, as thoughtful as Anthropic’s leaders are, they have failed to understand the only thing they can control is how they act. They can’t stop rivals from destroying the world, but they can decide they don’t want to be part of that.
They don’t have to make this AI monster. Instead, they are prioritizing being first above all else. Write to Tim Higgins at tim.higgins@wsj.com