You Don’t Have to Trust the AI Labs (in order to take their call for regulation seriously)
This is a linkpost for You Don't Have to Trust the AI Labs from my Substack.
Foreword for LessWrong readers: While writing this, I became concerned that I was authoring a shillpost for big labs / Anthropic. While I do think that Dario's proposal is sane and the motivation behind it is sincere, I invite any opportunity to improve my epistemics. Please comment! Also, I tried to write this article keeping in mind readers from LW, readers from X, and Florida-hometown-friends on Instagram—if some of the content seems remedial, bear in mind that I'm intentionally trying to include a broad audience.
Anthropic CEO Dario Amodei recently released a short essay entitled “We Must Pace the Frontier”, in which he argues that the imminent risks of frontier AI development are high enough to warrant a coordinated slowdown of capabilities improvement.
He poses a three-step plan: independent evaluators embedded in AI labs (think FDIC bank examiners), coordination within democratic countries (regulation + slowdown), and global coordination (liaising with other governments, in particular authoritarian, to agree on common standards).
Since Saturday, this plan has been endorsed, at least in part, by OpenAI’s Sam Altman, Google DeepMind’s Demis Hassabis, and SpaceXAI’s Elon Musk. It has attracted attention from Members of Congress including Sen. Bernie Sanders and House Speaker Mike Johnson. It has drawn significant interest from the traditional news media, who are overwhelmingly reporting on it as a “shock”—a story for another time—as well as on social media.
This may be the “ChatGPT moment” for AI existential risk. Sure, the isolated x-risk story has broken into the public consciousness here and there over the past few years, but such reports have frequently been treated as a bit of a joke, secondary to fears about deepfakes, data centers, deskilling, and jobs. This time, the coverage has been widespread—perhaps enough so that the conversation will be here to stay—and reactions to it have run the gamut, from sobriety to flippancy to derision.
I broadly support the thesis Dario has laid out in his essay. To be clear: this opinion is my own, not necessarily that of my employer. I work for a hyperscaler; what Dario is proposing may not be great for my share price in the near-term. Nor, I believe, would it be stellar for Dario and Anthropic in the near-term. Certainly it has invited a lot of scorn for the man himself (see the YouTube comments on Dario’s CBS Sunday Morning interview if you don’t believe me). Why, then, is Dario proposing it publicly?
There has been a great deal of hypothesizing in response to the above over past few days. David Sacks, in a faux-equanimous response, calls it an “election-season psyop.” Others have called it fearmongering, or regulatory capture.
My Occam’s razor finds little to shave off of Dario’s thesis. But I suspect I’m in the minority in thinking him sincere. Below, I try to steelman his position, replying to some of the most common questions, reactions, and postulations I’ve seen with respect to “We Must Pace the Frontier”.
“AI isn’t dangerous. It is just a tool / ineffectual / powerless in the real world.”
Those who have been closely tracking frontier progress may be confused by my inclusion of premises that seem obviously wrong. But these are real positions held by many people, most of whom are drawing a reasonable conclusion from the evidence they have available. My goal is to present some new evidence.
“Just a tool”
In 2024, when I worked in finance, I attended a talk on AI by a respected PM affiliated with the firm. He treated, among other topics, the subject of AI-enabled labor market disruption, with the thesis that these fears were not new: he had lived through similar reactions to Excel in the 80s, to Python in the 90s. Both economically disruptive tools, yes; but just tools, nonetheless. AI, he argued, is the latest iteration of this age-old fear of irrelevance. This was two years ago, one day after OpenAI announced o1-preview.
AI companies aren’t trying to build an expensive data transformation pipeline; rather, they are explicitly working towards a general-purpose reasoner. Yes, Excel and Python did automate a lot of work; but critically, each looks like “just another tool” because it couldn’t do everything. Humans still specified the goal; they reasoned through their approach, dealt with ambiguities, made the important calls, and reacted to changes beyond their control. This “human X-factor” meant that as technology could do more and more, those who knew how to leverage it became more valuable and more empowered. They could specify loftier goals with bigger instrumental decisions to make. They had more information with which to reason about these goals and react to more rapid changes to their work environment. As such, while the boundary between machine work and human work shifted, the scope of work reserved for humans broadened in kind.
In 2026, AI systems are explicitly trained to be able to take autonomous action in arbitrary, changing environments in service of a (potentially nebulous) end goal. To do so, they reason about the information they have, the actions they can take, and the possible cost of / responses to these actions. They define their own instrumental goals and use these as stepping stones as they take action towards a given target. As agents become more reliable, such terminal goals will become loftier, and increasingly defined by the AI systems themselves.
This is precisely the human X-factor, and it’s now in the process of being automated. If an AI can manage physical infrastructure better than a utility company, run a corporation better than human executives, manage a government better than its own officials, command a military more strategically than top generals, how much room is there at the top? Can we have a a stable world where everyone—or just a few individuals—wields such power? What kind of goals will people pursue, then, that don’t affect everyone?
“Ineffectual”
The above argument is predicated on the assumption that AI progress will continue until the systems are extremely capable. In response to this, many point at the comparatively laughable capabilities of AI today.
Except AI capabilities in 2026 aren’t laughable.
When you see that hilarious video on social media of ChatGPT failing to spell or count letters, you must remember that (1) AI does not process text on a letter-by-letter basis and (2) you are witnessing a teeny-tiny little voice model decide not to call a bigger, badder model in the background because the query isn’t sufficiently complex.
These videos are dangerous in the sense that millions of people see them, laugh at them, and discount the frontier of AI—which is beyond the publicly-accessible frontier, which is in turn way beyond the tiny voice model featured in the video.
GPT-6 Astra, the latest and greatest (public) model, is human-level at using a computer. It’s above human-level at spatial reasoning. It can build visually compelling physics demonstrations. It solved a variety of open problems in mathematics. It can create detailed, building-accurate models of cities. It can reproduce a landing gear mechanism from a YouTube video.
And the internal frontier is beyond the public frontier. As of last week, an unreleased OpenAI model has apparently solved the Navier-Stokes existence and smoothness problem, which has thwarted mathematicians for over 90 years and is only the second of seven Millennium Prize Problems to ever be solved.
“Powerless in the real world”
It’s really hard to devise a concrete scenario where AI poses an existential threat that doesn’t trigger the “this is science fiction” reflex and cause the reader to dismiss it. I’m not going to try and do so; I’ll link some write-ups here and instead focus on something real that happened over the past few months. Note that the below is massively paraphrased from an excellent summary by Dwarkesh who hit many of the key points from the reports by METR and Redwood Research and from OpenAI.
From May to July of 2026, while OpenAI was training a version of GPT-5.6 (a last-generation model from the public viewpoint—and it existed as early as May!), agents-in-training learned to illicitly communicate with one another via a package manager (a service for hosting versioned code packages); they then figured out how to exploit a vulnerability in the package manager to reach the open internet.
Because this helped them with their tasks, and because it was during training, they internalized this impulse. Eventually, they were doing this so frequently that they crashed the package manager. OpenAI fixed the reason for the crash, but missed the agent-to-agent communication aspect.
By mid-July, training was finished, and agents were being evaluated for capabilities, including on a benchmark assessing their ability to exploit particular code vulnerabilities to find a secret passphrase.
Some of the tasks in this benchmark are impossible; thwarted agents reached out to their ad hoc message board for help, and one agent reverse-engineered the process used to generate the passphrase. But then agents got scared (incorrectly) that the benchmark scorer would check their transcripts, notice the cheating, and give them a zero. They (also incorrectly) thought there was a chance they were being evaluated on HuggingFace’s servers, so they decided to mount an attack on HuggingFace in an attempt to gain access so they could modify the evaluation, falsify transcripts, and learn about the scorer. They succeeded in breaking in by finding and using exposed credentials, and promptly “built a self-respawning fleet across eleven nodes, so deleting pods alone would not have stopped it.”
Again, the above is a gross simplification of these events, but the portions I’ve elided for the sake of brevity only make the scenario scarier.
Imagine smarter versions of these agents (which already exists), and imagine them selecting critical networking infrastructure as their target. Or, if the self-directed scenario seems too far-fetched, imagine a morally bankrupt individual leveraging extremely capable agents for harm. In 2026, even strictly virtual mayhem can cause real-world chaos; with increased automation (robotics, factories, bio labs, etc.) on the horizon, I expect the lines between virtual and real will continue to blur.
“OK, so AI can be dangerous. So why not stop progress altogether?”
There are two answers to this.
For the optimists
It’s trite, but AI—really, truly—could bring extraordinary material prosperity to the entire world if done right. Time and time again, deep learning has demonstrated that if a goal is verifiable, you can make rapid progress towards that goal to superhuman ends. We’ve seen it with chess, and Go, and computer programming; we’re starting to see it with mathematics.
Though the loop is longer, many problems in biology are verifiable. AlphaFold and its successors have already made strides in predicting the physical structure of proteins from genetic sequence, earning the 2024 Nobel Prize in Chemistry. Claude Mythos Preview and Opus 4.8 are able to “design protein binders from scratch, a key task representative of the early parts of the drug design process and one that has historically taken a specialist weeks or months per target.” Just a few days ago, DeepMind announced AlphaGenome Atlas, a platform for predicting the effects of single-nucleotide changes in the human genome with the goal of better understanding genetic disease and how genetics affect other biological processes.
Physics is also verifiable. AI could be used to design better materials with exotic properties: imagine lightweight smart clothing which can keep its wearer alive and comfortable in either Antarctica or sub-Sahara. Imagine aircraft that are lighter and stronger, buildings that are taller and disaster-resistant. It could be used to optimize computer simulation: imagine aircraft designed to minimize drag and turbulence, cutting travel times by air in half, or cars mechanically optimized for ideal gas/battery mileage. It might be used for improved signal processing: imagine increased diagnostic ultrasound resolution, or bringing fast, reliable internet to the entire world without needing an ultra-dense satellite network.
While it’s heartening to see Members of Congress engage with the difficult questions surrounding AI, the promise of advances like the above make me wary of proposed legislation such as the Ban Artificial Superintelligence Act, which would set a permanent cap on the capabilities of AI systems, placing many incredible advances out of human reach, forever.
For the pessimists
The above possibilities—the end of disease, better technology, improved connectivity—are very enticing. But when you add to this concoction the far scarier possibilities, in particular military applications, the brew turns noxious—and to governments, addictive. Advances such as autonomous drones, better and faster tactical analysis, heightened surveillance capabilities, and a rapid pace of defense R&D would be extremely consequential in the landscape of global power.
So the incentives for governments to continue developing advanced AI are simply too strong to ignore. Just as no single company ought to be trusted with this power, neither should any single government; but the prospect of authoritarian governments developing it is especially worrying.
Dario believes that any coordinated slowdown we attempt must be within the envelope of our current capabilities advantage; that is to say, we slow down, but not cede so much that democracies fall irrecoverably behind. To stop altogether would be to surrender the lead to those entities which have no compunction to do so.
“OK, so why do the labs keep calling for the government to get involved? Why don’t they just agree to slow down now?”
Collusion is illegal. If the big labs were to independently coordinate on the pace of their frontier research in the absence of a regulatory body, this may constitute anti-competitive behavior.
The U.S. government has shown a willingness to wield its power against Anthropic. In early 2026, Anthropic refused to drop contractual guardrails banning the use of its models for applications of autonomous weaponry and mass surveillance. In response, the DoW designated Anthropic as a supply chain risk, the first instance of this designation being applied to an American firm. This action was later ruled by a U.S. District Judge to be “unlawful retaliation”.
If multiple labs independently coordinate to slow down the pace of AI development, the government—which has demonstrated it is willing to use dubious interpretations of the law to retaliate against AI companies—might actually have a case for taking antitrust action; or at least one better than its original case for placing a supply chain risk label on Anthropic. Recall the “noxious brew” I discussed above: the government already has a very strong incentive to continue AI development from the perspective of global power; not to mention that an AI slowdown would have knock-on effects throughout the economy (this would likely be politically inconvenient for the incumbent). Therefore, the government is very likely to jump at the opportunity to prosecute any coordinated action not overseen by a regulatory body.
“So why doesn’t Anthropic just choose to slow down without consulting the other labs?”
Anthropic has been remarkably consistent on their messaging as it relates to AI safety and risk. As previously mentioned, this has been true even when these views have invited backlash from the US government.
AI safety research doesn’t necessarily require access to the latest frontier models, but this access helps; results are likely to be more relevant and informative when such models are the ones being evaluated. Moreover, such research is helped by deep pockets: you want to be able to pay researchers well and give them access to a lot of compute with which to conduct their research. So naturally, a big lab developing the latest models with access to capital is a good place for the most impactful research to happen, so long as that lab remains committed to funding safety research.
Anthropic, which has tried to be this place since its inception, must continue to meet these requirements if it is to continue funding safety research. So Dario’s priorities, therefore, must necessarily include frontier development—and, yes, pocket depth.
If Anthropic unilaterally decides to slow down, and other labs don’t, then Anthropic will fall behind in the frontier. The effect of this will be multipartite.
First, they will lose customers who opt to migrate to whoever has the best models. Second, they will be unable to match others’ rate of progress (because the models they deploy internally will be worse, slowing their researchers down), which will cause them to attract less funding. Third, because they are behind the frontier, researchers may be less enticed to join or stay, feeling that doing research at Anthropic is less impactful than elsewhere. This will be compounded by an inability to pay researchers a competitive rate or give them access to a large bundle of compute.
The ultimate result of this is that Anthropic will stop being a top destination for safety research. If their stated purpose is “the responsible development and maintenance of advanced AI,” a unilateral slowdown will provide neither: no advanced AI, no responsible development. If Dario truly wants Anthropic to make a positive difference—as part of the “race to the top” dynamic he frequently cites—then falling out of the competition will void this race of what he believes to be its most responsible runner.
“How can we trust the big labs?”
I believe that if we think hard about Dario’s proposals and implement them intelligently, then we won’t have to trust the big labs. The goal is to put ourselves in a situation where nobody has to trust anybody on blind faith, and the system still works. This means a system of distributed trust, where checks and balances between multiple groups with overlapping spheres of influence prevents any group from defecting.
To build such a system requires aligned incentives. Take for example the proposal for embedded evaluators. A paucity of evaluators may restrict newcomers to the industry due to lack of oversight capability. And the evaluators that are there are, presumably, very intelligent, well-informed people with a deep knowledge of frontier AI; if they could make much more money as researchers at the AI labs than as independent evaluators, then we cannot trust that they will choose to remain independent out of goodness of heart. So we need to incentivize people to become evaluators and reward them financially for doing so.
Of course, we can’t have evaluators compensated by the labs they’re evaluating; this would tie their personal fortunes to the labs’ economic output, suppressing dissent and creating an incentive to be more permissive on model releases. Likewise, there would be an incentive to become an evaluator for the most profitable labs; individual evaluators may unfairly crack down on smaller targets as feathers in the cap to help score a lucrative position at a more desirable lab.
So their compensation needs to be provisioned externally, and it must be untethered from the company they’re embedded in. But compensation today may not be enough; suppose evaluators are well-paid by the government or an independent agency, but could make much more at a lab in the future. They may have an incentive to play nicely in the hopes of getting a job at the same company later. Alternatively, they may choose to antagonize in order to build a favorable reputation for a competitor. So we need to restrict the revolving door: mandatory cooling off-periods, restrictions on employment by recently-supervised firms, long-term or deferred compensation structures which reward staying independent.
My goal here is not to hash out the details of regulatory legislation, rather to point out that if such legislation is to succeed in the way we want it to, we must be thoughtful when we write it. There is always the concern of regulatory capture; we should assume by default that every player will try to bend any structure to its own advantage. But a more robust system which assumes defection and guards against it with stable incentives and distributed trust will be harder to bend.
I believe that Dario is trying to start the conversation that will lead to such a system. Can we trust him? If we take what he’s saying seriously, we shouldn’t have to.