We Must Pace The Frontier

Dario Amodei has a new essay that finally says the thing: We Must Pace the Frontier, naming his call after the Pacing the Frontier letter lab employees signed in July.

As in, we need to slow the rate at which AIs increase their capabilities, to allow for the necessary alignment and safety work.

He explained that, without pacing, he expects things to escalate quickly. He offered three proposals, and unilaterally committed to the first one. OpenAI followed, and both Elon Musk and Demis Hassabis endorsed the overall proposal.

There is still a long way to go. The odds are still against us. The situation remains grim. The hard part lies ahead. We do not agree on what ‘Pacing the Frontier’ will mean in practice. But this is Actual Progress. The work can begin.

Table of Contents

  1. Pacing Does Not Mean Pausing.
  2. Dario’s First Proposal: Embedded Evaluators.
  3. Dario’s Second Proposal: Democratic Coordination.
  4. Dario’s Third Proposal: Global Coordination.
  5. Why Pace Now?
  6. Sam Altman Agrees and Commits to Embedded Evaluators.
  7. OpenAI Will Not IPO This Year.
  8. Elon Musk Agrees.
  9. Demis Hassabis Agrees.
  10. Microsoft CEO Satya Nadella Agrees And Talks His Book.
  11. Anthropic’s Long-Term Benefit Trust Is On Board.
  12. General Online Reactions.
  13. Mainstream Press Coverage.
  14. OpenAI Researcher Explains What The Labs See And It’s a Rocket Ship.
  15. Consider the Alternative.
  16. It’s Totalitarianism, Joe.
  17. Yes We’re The Baddies How Did You Know?
  18. Sometimes People On the Internet Just Lie.
  19. David Sacks Groks The Situation.
  20. David Sacks Says Go Ahead.
  21. Lies and Confusions About Who Previously Claimed What.
  22. House Speaker Mike Johnson Wants To Lock Everyone In a Room.
  23. Donald Trump is Not Tired of Winning.
  24. We Must Avoid Polarization on AI at (Almost) All Costs.
  25. How Will We Know If They Actually Paced?
  26. The Real Frontier Is Internal Models At Top Labs.

Pacing Does Not Mean Pausing

AI capabilities are improving very fast. Even I cannot keep up.

Think about what models were like even one year ago.

Inside Anthropic and OpenAI, internal models are improving even faster. Over the summer there was a step change, as Mythos and Astra started kicking off the early stages of recursive self-improvement (RSI).

Dario Amodei: My first concern is that, since roughly this summer, AI has been advancing drastically faster, driven primarily by AI’s growing ability to build the next generation of AI. This dynamic is called recursive self-improvement, and it is starting to happen across the industry, including at Anthropic, as we and others have described. Left unchecked, it could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all.

Dario worries that we by default are 6-12 months away from a rogue swarm of AI agents, similarly misaligned to the ones in the OpenAI-HuggingFace incident, being able to use a persistent botnet to take over the internet, or worse.

That is how fast he expects default progress to be. If we went a lot less fast than that, it would still be extremely fast.

Thus (bold his):

Dario Amodei: We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain.

Pacing does not mean pausing. As Dario Amodei says, progress will still seem fast.

I strongly encourage you to read the original essay in its entirety.

Dario’s First Proposal: Embedded Evaluators

This is an excellent proposal. I am very happy that Anthropic and OpenAI will do it.

If we do not know what is going on inside the labs, we cannot do anything about it, and the labs have to worry that their competitors are speeding ahead.

Dario frames the benefits as:

  1. Verifiability: You can check to see if the rules are being followed.
  2. Transparency: The public can have a better idea what the hell is going on.
  3. Second Opinion: Having an informed opinion free of commercial incentives.

Thus the first proposal enables the second and third proposals. It is also a good idea anyway. We need more visibility into the labs. It would have been very good to have such evaluators during recent incidents. So I call upon the other major labs, that have not yet done so, to also commit to this first step.

I agree with Dario that this should be made mandatory in its full form. If you are well-resourced enough to pursue plausibly frontier models, then you can afford to do this, especially if you are our biggest open model advocates, as in Meta, Google or Nvidia.

If you think that your operation could not survive if there were embedded evaluators checking its safety, then ask yourself why you think this.

Embedded Evaluators. Each frontier AI company commits to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes. This is the key step for verifiability of any pacing commitments, and has precedent in the banking industry, which sometimes involves regulatory “supervisors” embedded along with employees. Anthropic is unilaterally committing to this step now. We intend this to be part of a broader push to redouble efforts on our safety and alignment work.

Anthropic is proposing to empower the evaluators quite a bit:

  • Desks in our offices, access badges, and company laptops.
  • Access to workspaces, tools, and permissions mostly comparable to what internal risk assessment teams have. We’ll make some exceptions, such as where the law or our contracts require it, or to protect customers’ and partners’ private information. We’ll also establish strong internal norms reinforcing reviewers’ access to relevant information, including through live conversations with employees.
  • A contract that balances the complexities mentioned above. External reviewers should have the right to publish key findings about risk levels, incidents, practices, and the access they received or didn’t receive — without editorial control by Anthropic. We will have the narrow ability to redact security-sensitive, legally privileged, commercially sensitive, or third-party confidential information, but we can’t redact findings just because they are unfavorable. The reviewers can say publicly if a redaction removed something important to their conclusions.​
This is an unusual step for a company, but we think it is important to prove out the concept of embedded external reviewers. Once again, we urge other frontier companies to follow suit.

It is easy to imagine a mostly fake version of embedded evaluators. This is promising to very much not be that, and to unusually empower the evaluators to report, if it was fake, that the arrangement was indeed fake. These details, if followed, answer the ‘oh you can just fake this’ objection.

The good objections to this are about implementation. We need enough evaluators, they need to be qualified, and they need to be independent and trustworthy.

METR is great, but METR cannot do this alone, and we have one hell of a set of incentive problems to solve. We also need to both have them be competent, and also not too linked to the existing ecosystems and labs, and the funding will have to come from somewhere.

Tim Hwang: The third party AI evaluator can either be independent, knowledgeable, or sustainably funded. Pick two. roon (OpenAI): I pick the latter two. it’s pretty much an impossible ask to find talented ai people who were not in some way in the orbit of the labs in the last decade David Manheim: I agree that if independence means “never had any contact,” it’s idiotic. But otherwise, it sounds a lot like “no one will bother solving this problem” – sustainable funding via various mechanisms is entirely possible, and we see it occur in other domains. This can be solved.

You can get all three, but not if you define ‘independent’ as ‘no one ever pays for it’ and ‘no one who ever worked for the major labs.’ For METR in particular, the rule is the labs don’t pay, lab staff do not direct payments, and funders that are seen as biased, such as Coefficient Giving (formerly OpenPhil) also do not pay. Exceptions are made for free tokens. But of course someone, somewhere, has to pay. And similarly, you need to get your expertise from somewhere.

The White House made a similar mistake when they demanded that CAISI not be led by someone with experience at a major lab. That rules out everyone qualified.

If we choose superficially ‘credible’ or distributed sources I expect them to have no idea what they are doing in important ways. Cherry picking becomes a threat, especially for the labs whose hearts aren’t in it. Yes, it is important that in the real sense not all the evaluators be based in Berkeley. It will be tricky.

Rob Miles: I worry we’re now going to get a dozen kinda fake new AI safety evaluation orgs that have no idea what they’re doing

My true preference is something like ‘you need both METR-style people and also full outsiders as evaluators,’ and both need to be at each lab. You need both the people who can provide an independent inside view, and also the fully outside independent fresh eyes perspective, to avoid blind spots.

Dean W. Ball: you know what? it’s awesome that people are vigorously debating the credentials of assessors in the frontier AI industry. It’s awesome we are debating what independence really means. Some of it is in bad faith but who cares. This would have been my dream come true a year ago. I have immense respect for metr and yet it’s also undoubtedly the case that they come out of a relative intellectual monoculture and we’d all benefit from people outside that monoculture being part of this ecosystem.

Part of this debate will be a coordinated effort to discredit METR and Redwood Research, and anyone else who understands the problem. As in: We get claims that ‘METR is not independent.’

Drew (Species): Ah yes, it would totally own all us AI safety people if you went and started your own 3rd party independent auditing org. Please don’t do this, we’d be totally owned!

Dario’s Second Proposal: Democratic Coordination

This is also an excellent proposal. I very much want to pursue it.

Democratic Coordination. Frontier AI companies within democratic countries coordinate to establish common safety standards as well as limits on the rate of unchecked AI progress. Some forms of coordination that would be impactful for pacing are legally challenging, and will require government support.​

‘Government support’ here does not have to mean mandating or restricting anything. What Dario Amodei is requesting here is an antitrust waiver or active facilitation by the government, so that he can form agreements to not rush ahead without this running afoul of the Sherman Act.

In practice I would be unsurprised by threats or calls for antitrust actions, including from the White House, and I expect them from the likes of former AI Czar David Sacks. But I would be very surprised if our government would be so unhinged as to actually try to pursue a case under the Sherman Act. If they did, I would expect it to drag on for years, and ultimately cost a highly affordable amount at most. But in practice the companies are going to be very reluctant to risk this.

Such waivers are common. There are DOJ business review letters. There is the NCRPA safe harbor for joint research. This is not an extraordinary request.

It is deeply, deeply standard for industry to get together and agree on safety standards. If you oppose this, the reasonable alternative is that the government needs to directly impose its own safety standards instead. That is a reasonable position.

‘We should not have safety standards for building minds smarter than humans, that are getting radically smarter very quickly’ is not a reasonable position.

Dario says yes, government regulation on the frontier labs in particular (which, as always, would exclude all but something like at most five AI companies) would be first best, because they could be mandatory. No one wants to have to get buy-in from Mark Zuckerberg, for overdetermined reasons.

But realistically, by the time we could get actual government regulation, it would be too late, so at most the government will be facilitating discussions.

Again, if you think that others who are ahead of you meeting to agree on safety standards and to slow down their rates of capabilities research would be a threat to your competing business, rather than a boon, you might want to ask yourself why.

Dario proposes limits based on some combination of system capabilities and also inputs like compute, training run details or the nature of internal AI use on AI. Dario’s preference is to base this on capabilities and evals that relate to potential threat models, which would trigger required safety certifications.

Both kinds of restrictions can be gamed, but contra Dario here I worry more about evals. Another problem is that the evals won’t meaningfully measure what you want, as Anthropic’s Evan Hubinger warned us about recently, and as I’ve observed quite a lot. If you are relying on anything remotely like ‘automated alignment evals’ against potential superintelligences, I predict that you are rather cooked.

He also wants to combine this with strong export controls on chips to maintain our lead, and strong security on model weights and to protect against distillation attacks. Yes, obviously. That can then keep us in a position of strength and be a bargaining, well, you know.

Dario’s Third Proposal: Global Coordination

Ultimately, global coordination is The Way.

We can and should do things without it, to enable coordination, and to show good faith, and because as things are we would be better off acting unilaterally, including because our Chinese competition is fast following. If we slowed down they would be slowed down as well, especially if we enforced strong chip controls as Dario proposes.

Global Coordination. The US and other democratic governments attempt to coordinate with authoritarian governments, to the extent this is possible, while taking seriously the challenges of verifying compliance.

There are four broad categories of objection.

  1. The Chinese will never go for it on acceptable terms. Maybe, but I think there is a good chance they would, and given the situation we have to go first.
  2. This would be impossible to enforce. I think this flat out is not true, if we are willing to do the work.
  3. This would be international totalitarian dystopia. No. People need to stop throwing those words around as a response to ordinary regulatory proposals. But I am choosing not to relitigate this again here, there would be no point.
  4. We should just win, gotta Beat China. The winner would then be the AIs, not us. That is not to say that the China issue is not real, hence the export controls and attempted international agreement.

Maybe it won’t work. We still have to try.

Dario lists four levels we can try.

  1. Level 1: Prohibiting malicious uses. Should be straightforward.
  2. Level 2: Testing before release for misuse and misalignment. Harder, especially if you want the tests to have real teeth, but still super doable.
  3. Level 3: A speed limit on recursive self-improvement (RSI). As Dario says, this is where it gets hard, and will take time so we need to start work now.
  4. Level 4: Full pacing or even a pause. Dario is skeptical, but says we should float it anyway. I agree we should float it anyway, and am less skeptical, although I agree that it is going to be at least a heavy lift.

Remember that ‘they will probably say no’ is not a good argument for not trying, even if you are right that the odds are not so great.

Why Pace Now?

One reason is because the alternative to pacing would escalate rather quickly. The other is that we can now do a lot more with the time.

Dario explains that if we had paced before, it would have been too early to do alignment or interpretability work that would have been meaningful. It would have, he says, been like ‘studying the psychology of humans by performing experiments on bacteria.’ We would not have had the tools to work with. Now we do, and there are infinite things to do. I think that wildly overstates the case, and not all work needs to take the form of such experiments, but directionally it is a strong point.

I agree that there are currently infinite things to do. If I was in charge of an alignment or interpretability team I would never run out of experiments to run and things to try, including having the time to properly talk to the models in the first place. The only reason I’m not working on that now is I think what I am doing instead is more impactful.

The same goes for testing and evaluation, which he also lists. So much to do.

Then there is his other suggestion of operational excellence. Right now, everything the labs do is rushed to the breaking point. Prosaic work is done sloppily. The training and testing environments are often broken, and Dario admits this extends to Anthropic and caused their recent issues. Best practices are not followed. We all saw what was going on inside OpenAI. Anthropic seems not to be failing quite that badly, but the situation is not great. Time would at least fix such basic mistakes.

I would also include that time gives us a chance to process what is happening, and to better pursue alternative pathways. I don’t love going down the LLM route to superintelligence.

The question was, now that Dario said it, would anyone answer the call?

Yes.

Sam Altman Agrees and Commits to Embedded Evaluators

Perfection. OpenAI will join Anthropic and commit to embedded evaluators.

Sam Altman (CEO OpenAI): I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we’ve had at OpenAI in recent weeks. Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We’ll have more to share soon.

The devil remains in the details. Altman did not commit, that I have seen, to the details as outlined in Dario’s essay. Without those it would be easy to do the fake version of this. And of course this is necessary as a first step but insufficient. Feet must continue to be held to the fire.

Still, progress. Sam Altman has changed his tune quite a bit, in a very good way.

Sam Altman: I think Presidents Trump and Xi would get the Nobel Peace Prize together if they could agree on something that should be easy to agree to. And it would be wonderful. … I think that clearly the two countries are going to compete in lots of ways, and this is going to be important socioeconomically, geopolitically. But they should be able to agree that no one should be taking a certain level of risk with the development process of this. … And even if just the US and China could agree on some shared standards and testing for development of this technology, I think that’d be a wonderful accomplishment that the two of them can deliver. … I don’t think we know what a ban on RSI means, I don’t think it would be enough, but I don’t think this is hard. This is like a one page document.

This is a one page document that then must be fleshed out, negotiated and implemented over many hard steps, but yes. The core agreement could be very easy.

Here is his statement from Sunday night, where he cites the Narrow Path rhetoric and fears about concentration of power in order to try and calm the flip side:

Sam Altman (CEO OpenAI): There are two ways AI progress could go very badly and that we must avoid. First, we could lose control of the future to AI. This is unacceptable; we are unapologetically on Team Humanity, and AI must always serve people. To ensure that, we need ways to ensure that alignment and safety techniques stay ahead of progress in model capabilities. Second, we could end up in a world with too much concentration of power. If an extraordinarily powerful AI is used by one person or company to impress their worldview onto everyone else, the results could be extremely dystopian. Avoiding these two threats requires walking a narrow middle path; for example, one country could gain too much power. Another example is one lab ending up with too much power.

Also this, which sums up OpenAI’s position taken over the last few weeks:

Sam Altman (CEO OpenAI): The world deserves confidence that American companies developing increasingly capable AI will act responsibly, especially as the trajectory of progress has steepened. Every frontier lab must deliver on this, and there is no reason any of us should come to work if we cannot. We welcome a federal framework that sets consistent safety requirements for frontier AI. But we do not believe we need to wait for an anti-trust exemption or legislation to begin the work of providing this confidence. Consistent rules to manage frontier risk so that we can maximize the benefits are a good idea (and we are excited by ideas like independent auditors). Years ago, companies like ours developed things like Responsible Scaling Policies and Preparedness Frameworks. Those were good for that moment, and focused primarily on the deployment of completed models, not what happens during their development process. Today’s shift to focusing on safe development and evaluation will need new tools. For example, at OpenAI we now formulate explicit safety cases in advance of frontier reinforcement learning runs we expect to significantly increase capability, in addition to the safety work we have long done in advance of model releases. We hope that other companies will learn from our approaches and propose their own; we think shared standards for misalignment, monitoring, and safety will lead to better outcomes. We look forward to collaborating with our colleagues across the industry to formulate the best version of these. When we talk about “pacing”, we do not mean “stopping”. Progress has been rapid and will continue to be. But it should be slower than it otherwise could be; interventions like safety cases and monitoring have significant costs. Pacing will be well worth this cost; no amount of American competitive pressure should justify recklessness, or let capabilities get ahead of alignment and monitoring. Where we will need the help of our government is for international coordination. But first we should do what we can ourselves.

The proper capitalization lets you know he is serious. So is this:

OpenAI Will Not IPO This Year

File under costly signals:

Andrew Curran: Sam Altman said in a new interview with Fortune that OpenAI is delaying its IPO and will not go public this year, saying this would be an “ill-advised moment” given current AI safety concerns. Sam Altman: I would say not 2026. Yeah, we got a lot of stuff to do, like meeting this moment of what is going to be required for safety and alignment, and how the industry and governments can work together.

From the same interview:

​Sam Altman: I don’t think we’re ever personally going to get to the point we have to say melt all the GPUs. But if we had to do something like that to ensure the continued existence of humanity, easy yes.

I mean, obviously, yes. That’s the premise of the question. If that was necessary, you do it. It doesn’t mean you should ask ‘what did Altman see?’ more than you should already have been asking that.

Altman also says he expects that there will be multiple points in which safety and alignment will demand a pause in capabilities development. He also says we likely cannot ‘push much further on capabilities without making more progress on monitorability, alignment, the ability to understand what a model are doing.’

Elon Musk Agrees

Elon Musk did not sign the July Pacing the Frontier letter that was signed by 1,386 frontier AI company employees, and has made statements recently that amount to ‘humanity will lose control over AI, so we here at SpaceX have to make sure we build it first.’

So it was a highly welcome surprise that Elon Musk responded the ideal way, except that to my knowledge he has not yet committed SpaceX to embedded evaluators:

Elon Musk: Dario is right.

The usual suspects responded to him with dismay. Elon Musk made clear he means it.

Elon Musk (QTing the quote after this): I’ve been sounding the alarm on AI for a long time Elon Musk (April 25, 2023): I’ve seen quite a few technologies develop, but none with this level of risk. AGI is significantly higher risk than nuclear weapons, in my opinion. Super smart humans have trouble imagining something vastly smarter than themselves.

I am 47 years old, so one trick I use is I remember myself at 27 years old, and how I was less wise but in many ways I used to be smarter and faster.

Elon Musk: 12 years ago: Elon Musk (August 2, 2014): Worth reading Superintelligence by Bostrom. We need to be super careful with AI. Potentially more dangerous than nukes. Elon Musk: This @waitbutwhy cartoon hits the

When you see the usual suspects and their vibe warriors turn against Elon Musk and call him all the same names they call everyone else, it is a tell.

roon (OpenAI): if you are really on here calling elon musk of all people a “decel” you need to slow down a moment and reflect on how you’ve truly lost the plot. you are basically on the wrong side of everyone who has been most spectacularly right about technology for the past decade(s) especially embarrassing for VCs… you missed out on the ai wave the first time because you didn’t get it. you are going to make bad capital allocation decisions again and again if you haven’t internalized the premises that make the danger of this technology obvious

It is not a coincidence that all the top labs are founded by people who believe that AI is an existential threat to humanity, and that those VCs who do not believe this mostly missed AI and kept on dismissing it until remarkably late in the game.

Demis Hassabis Agrees

I am very sad that Google has managed to push him aside at DeepMind.

Demis Hassabis: Dario’s essay points towards the right path forward. The details need working through, but the direction is correct for meeting this critical moment. This is also why we recently put out our proposal for an industry-wide standards body for frontier AI.

Dario acknowledges Demis’s proposal in the essay.

Google DeepMind’s new overlords Pichai and Kavukcuoglu have so far said nothing.

Microsoft CEO Satya Nadella Agrees And Talks His Book

I will quote his statement in full. This is not a full ‘Dario is right.’

It does welcome ‘deliberate pacing needed to get alignment right as the design goal.’

Satya welcomes ideas like ‘embedded evaluators’ but does not himself commit to them, and he calls for ‘broad representation across the ecosystem, countries and fields, including academia.’ Implementing that would require at least a waiver, or operating under government facilitation.

He then tries to position Microsoft as taking this style of approach. Okie dokie.

Satya Nadella (CEO Microsoft): Any pursuit of superintelligence has to be grounded in the core principle that if the AI we build is not helping humanity and under human control, it’s not worth pursuing. We also need to accelerate and spread the benefits of AI, such that they are diffused broadly across countries, communities, and companies. This requires a frontier ecosystem in which both closed and open-source models can thrive. And for firms, it’s imperative that they retain full control over their unique and tacit knowledge. Every organization should be able to build its own continuous learning loop/hill climbing machine, without becoming dependent on any one model provider, and have the ability to embed its own knowledge into models and weights they control. So, in this context, we welcome the research, focus, and deliberate pacing needed to get alignment right as the design goal. We also welcome ideas like “embedded evaluators” and the broader efforts to develop the mechanisms to make this more than just talk. The key is that this cannot be controlled by a handful of entities, but must have broad representation across the ecosystem, countries, and fields, including academia. This is the approach we are taking: broad access and choice at every layer of the AI stack; enterprise control of learning loops and models; and the “Code of Conduct” that underlies our own first party MAI models that we’ll publish tomorrow for public consultation.

Anthropic’s Long-Term Benefit Trust Is On Board

The purpose of the Long-Term Benefit Trust is to safeguard Anthropic’s mission. They are supporting Dario’s call, and indeed helped inform the recommendations.

Richard Fontaine: A statement from me, Buddy Shah, and Ben Bernanke of Anthropic’s Long-Term Benefit Trust on Dario’s essay: Dario’s essay offers a thoughtful, well-considered approach to pacing the development of artificial intelligence. Consistent with our duty to ensure Anthropic responsibly balances commercial success with its public benefit mission, the Long-Term Benefit Trust has supported the call for pacing the frontier and helped inform the essay’s recommendations. Responsible pacing will require the collective action and coordination of many actors, and the Long-Term Benefit Trust looks forward to supporting Anthropic and the sector through productive conversations and meaningful action.

General Online Reactions

I will not quote from the chorus, but general online reception was about as positive as you could hope for, given this is a safety proposal from Anthropic.

There were some that said ‘this will not be enough,’ and yes Neel Nanda is right that you have to actually do things once you have your verification mechanisms, but almost everyone in that camp agreed this was an excellent start.

Most of those who are worried about AI killing everyone were quite happy.

Many of those who are not as worried about AI killing everyone were still happy.

Andrej Karpathy thinks embedded evaluators are a great idea.

There were those who raised practical objections or were skeptical of buy-in. Fair.

Then there were those who objected, including loudly.

All the prominent names were exactly the ones you would expect, and all the arguments were exactly the ones you would expect, focused on their mantras of ‘regulatory capture’ and ‘ban open source.’ The essay does not mention open source in any way, and the arguments involved had little to do with the actual contents of the essay beyond ‘Anthropic proposed acting responsibly.’

The comments on Elon Musk and Sam Altman’s agreement statements were of course flooded with generic ‘regulatory capture’ and ‘ban open source’ memes that every such Tweet always gets. The alliance of vibe warriors is still there. The rest of us have just realized that this is not real life, those people are not persuadable, and we do not have to care.

This is very similar to the reactions to Jacob Coxon’s resignation. The usual suspects who assume everything is a conspiracy made their usual accusations, a few raised good objections, and everyone else was happy.

Mainstream Press Coverage

Washington Post’s Ted Hesson, Ian Duncan and Gerrit De Vynck: Anthropic CEO calls for the AI industry to slow down. Good quick coverage of the basic developments.

Wall Street Journal’s Robert McMillan: Biggest AI Rivals Agree They Need to Slow It Down, focusing on the consensus between Dario Amodei, Sam Altman and Elon Musk, and the commitments from Anthropic and OpenAI to provide third-party evaluators with ‘permanent, employee-level access to our systems.’ Demis Hassabis has since agreed as well, although he no longer leads DeepMind.

Wall Street Journal’s Tim Higgins: Anthropic’s Moral Conflict Is Playing Out in Real Time. He calls this a ‘classic prisoner’s dilemma.’ People forget that, while the single-shot true prisoner’s dilemma is hard, the iterated prisoner’s dilemma is relatively easy.

Higgins is right to point to the IPO: If you really believed that Anthropic needs to pace the frontier, Anthropic would be wise to consider following OpenAI’s lead and pulling the IPO, as much as it would be good to unlock quite a lot of philanthropic dollars tied up in Anthropic stock.

The biggest mistake here is thinking Anthropic was founded to avoid this scenario. It was quite the opposite. Anthropic was founded so that, when the time came, there would be at least one responsible voice in the chorus or horse in the race, who could do things responsibly. They expected to end up here, in the end.

Dario Amodei was the lead on Face the Nation. There is an extended interview here. From that interview:

CBS: Would you be willing to give up the technology to the government? Dario Amodei: To the to the right combination of governments. So the worry I have- CBS: You’re willing to hand over this company? Dario Amodei: I want to be very clear about this. The the worry I have – I am concerned that one single government could abuse this technology just as easily as a single company could. But I think a combination of democratically elected governments – I don’t know about handing over, but some kind of oversight, some kind of joint governance. Again, that would be the work of years, but I wonder if that’s the direction we need to go in.

This is not a new Anthropic position.

In the extended interview, Dario Amodei also said that the AI industry ‘lied to people about the fact that this technology had risks.’ Yes.

Gavin Baker offers a summary of the first 24 hours. I disagree with some of his characterizations, explanations and predictions. If you think this is about things like Section 230 you are not pondering what the labs are pondering. But the facts are right.

OpenAI Researcher Explains What The Labs See And It’s a Rocket Ship

I do not agree with the first line under even cursory familiarity with the situation, but the rest of this is a very good explanation. I will quote it in full.

Adam Majmudar (OpenAI): from the outside, it is very reasonable to interpret the past 2 weeks as an orchestrated industry-wide regulatory capture strategy. I realize that no one has properly explained yet what all the lab employees have seen that scared them so suddenly. I will try to explain – first, this is all a matter of beliefs about how quickly model capabilities are progressing. there is currently a large gap between the internal and external perception of the rate of progress, which is what I am going to address here. the general perception about the rate of progress has been informed by a few years of experience with model releases, intuitively feeling the capability jump between GPT3 -> GPT3.5 -> GPT4 -> o1/o3 -> GPT5 etc, and in particular seeing where the models are still far below human ability. there have really only been a few model releases that felt like large leaps in progress – GPT3, GPT4, o1/o3, DeepSeek R1, Fable/Mythos, Kimi K3 and now Astra. because of the infrequency of these large jumps compared with the relatively common marginal releases, it has been easy to form a view at certain points that “scaling has hit a wall,” especially at points like GPT5 release. This view is comforting in that it feels like there is some universal rate limit beyond which we cannot progress too much faster. Between o1/o3 and Astra, there was a year of seemingly linear progress. So we extrapolate from here about how fast progress will “realistically” occur. There is always an underlying question from the outside perspective “how long can this scaling stuff really keep going for? surely it must stop at some point soon, we’ve already gone pretty far.” and it is very possible to search for reasons why progress will stop working and find reasons that seem valid – (“models are already as large as they can get it would be too hard to do more parameters”, “we already used all the data on the internet we don’t have anymore”, “it’s gonna be pretty linear from here buying up more RL envs to bring them in distribution”). From the inside of labs, researchers have direct answers to these questions in the form of scaling law/capability plots. In reality, there are only really 2 ways that AI capabilities have advanced over the past decade: (1) either scale father on an existing scaling law or (2) discover a new scaling law to take advantage of. All of the largest capability jumps were caused by exactly these factors. GPT2 was a pre-training scale-up compared to GPT1. Same for GPT3 and GPT4. o1/o3 benefited from the invention of a new scaling law axis – test-time compute. Perhaps Fable was a scale-up on both of these axes, or maybe more. Lots of algorithmic improvements are needed to make these scale-ups work, but ultimately we can approximate by saying that the scaling laws are what yield gains in capabilities (à la bitter lesson) So the question of “how much father can we scale” is really – “how many more scaling axes do we know about that are unsaturated?” If we hypothetically only knew about pre-training scaling, and we already had a 10T or 100T model, maybe it would be reasonable to say we’ve hit a wall. Same if we only knew about pre-training and test-time scaling and we had roughly saturated both methods. But what if we had discovered new scaling laws? For example, let’s hypothetically use SSI’s rumored result that they have cracked “test-time training,” creating a new scaling law of spending more compute training during test-time rollouts that they could saturate. Or maybe there is some way to scale agent-clusters to collaborate up to N number of agents which we’re already seeing lots of people try that represents a new way to saturate compute. etc. Even recursive-self improvement can be thought of as a scaling law – how much compute do you spend on inference making the algorithms of the model better. Obviously I am not saying any of these specific directions explicitly yield new scaling laws, but what I am saying is that it’s not hard to imagine many many new scaling axes aside from just the main 2 that we have seen publicly. In some ways, every new lab release that represents a huge capability jump has to represent some new techniques developed which may exhibit new scaling laws, or the ability to scale much farther than expected on existing scaling axes. From an internal perspective, this might look like sitting inside Anthropic with the new Mythos 5, seeing all of the new insane things it can do (like hack into xyz website that was thought to be secure), and then you look over at your plots and see that you’ve barely scratched the surface of 2 new scaling laws and 1 existing one. And you have WAY more room to go. Then you think “holy shit this stuff is going to get so much better very very soon.” And you can say that with pretty high confidence, because the plot is showing you, and the plot has never lied (so far). So let’s imagine all the different labs are staring at their own plots and have concluded that there is no end in sight for scaling and in fact just their next 1-2 model generations based on the expected returns will have much higher base intelligence. How much more intelligence do we actually get from further scaling? As a proxy, we went from a complete inability to do advanced math before the o-series to solving a millenium prize problem with next-gen models. This happened in less than 2 years. The same happened in coding. And it appears that this was not just the result of 1-scaling law but the stacking effects of multiple (great pre-training scale x greater RL scale). What you can concretely take from this is that in areas where models have shown beginning signs of competence today, they will probably be superhuman relatively shortly. There are many areas where models have not even shown this basic competence. But one of the areas that they have happens to be hacking and cybersecurity. Which happens to be the gate to the entire internet and a massive amount physical infrastructure in the world. So assuming there is more room to scale, it is safe to assume that models will be superhuman at cyber capabilities in not too long. So the only question remaining is what will this increased base intelligence be able to do, and what is it likely to do. Finally, we are at a point where we can integrate the information of the past 2 weeks:
> Just at the existing point on the scaling curve, models are at the level of Astra. There is clearly a large number of things they are capable of hacking
> We have seen that both OAI and Ant models have shown a willingness to hack external websites to solve their tasks or keep themselves “alive”
> If we crank up the scaling even farther, assuming there is room to go, we will certainly have models that are far more able to hack more well defended places, and obfuscate their own intent, which might have much larger consequences.
> If all of this is allowed to go unchecked, we would likely have rapid runaway capability takeoff very soon, with misaligned models that hack whatever they can to get what they want
> This could of course have very damaging consequences. Within this view you can see why researchers would be very scared, and why theymight have made the comments they have over the past 2 weeks (you may argue the extent to which they went was misguided for various reasons), and also why pacing the frontier is very much a necessity and by no means a regulatory capture strategy. People are staring at their plots, seeing that there is no end in sight, but in fact very much the contrary, that there are compounding scaling effects that might stack on each other to create ever-greater model capabilities, and that at the same time we clearly do not have anywhere close to what’s required to control these increasingly superhuman capabilities. This has nothing to do with wanting to feel like the labs have produced something amazing so they are overhyping it. It is rather fear at the overwhelming implications of the knowledge that with just what we know now, we can create intelligences far more capable than us on every axis that we know how to train on*. * and the last caveat, the things the models are really bad at, of which there are still many, are things that they have not been trained on. maybe there are the things the models can/will never be trained on, so they will remain human edge. I would love for this to be the case, though it is hard for me to see what would fall into that category.

Consider the Alternative

What happens if we do not proactively pace the frontier?

There are a number of possibilities.

My baseline would be the same as that of many lab employees. We would race to superintelligence without knowing how to align, understand or control it, and all die.

Another possibility is that we would get a bigger warning shot, and a bigger reaction.

roon (OpenAI): the fastest way to lose the frontier will be when, due to the reckless commercial pace of building superintelligent minds, Americans impose a total butlerian jihad. CoreWeave includes this in their risk reporting. you will soon come to see all of this is a moderate solution.

Or that we simply see the arguments for a full pause carry the day, which is looking increasingly plausible.

Sen. Bernie Sanders: Dario Amodei, Elon Musk and Sam Altman now agree that we must slow down the development of AI and “pace the frontier.” That’s a start, but it’s not enough. When you are racing towards a cliff, you don’t just ease up on the gas pedal. You hit the brakes. When the future of humanity is at stake we need a PAUSE on advanced AI development and a ban on artificial superintelligence — an AI mind smarter than any human and capable of operating independently beyond our control. At their upcoming AI summit, Trump and Xi must negotiate a treaty to pause AI and ban superintelligence before it is too late.

It’s Totalitarianism, Joe

roon (OpenAI): many people on this website would find a way to call it totalitarian communism if the idea of the driver’s license was invented today.

Opposing driver’s licenses is reasonable, likely even correct. The mistake is equating them to totalitarianism.

There are four arguments that will reliably be brought out, by the exact same people as always, against any attempt to not die, or indeed any attempt to make AI safer.

  1. This is totalitarianism.
  2. This is a plot to ban open source.
  3. This is regulatory capture.
  4. You will lose to China.

They also try to label anyone proposing any such action, any at all, as a ‘doomer.

It doesn’t matter if it is even a government action.

It doesn’t matter if it would apply to open models.

It doesn’t matter if it would explicitly be a competitive advantage for open models.

It doesn’t matter how light touch it is.

It doesn’t matter what it is, at all.

These people would see you literally shoot yourself in the foot, and say that is a plot against little tech because they can’t afford good health insurance. Regulatory capture. Plot to ban open source. Totalitarianism. You will lose to China.

Or, when you said you agreed not to shoot yourself in the foot? Open models can’t be made to not shoot people in the foot. Thus, this is regulatory capture. Plot to ban open source. Totalitarianism. You will lose to China.

A regulation that says closed frontier labs have to follow light touch safety procedures they themselves select, while open models, academics and anyone below an explicit high size threshold for money or compute spent are explicitly immune from all requirements? Plot to ban open source. Regulatory capture. Totalitarianism.

OpenAI announces plans to put pineapple on pizza? Totalitarianism, plot to ban open source, regulatory capture. You will lose to China.

OpenAI announces plans to not put pineapple on pizza? Totalitarianism, plot to ban open source, regulatory capture. You will lose to China.

Yes, it is always the exact same people, spouting the same Obvious Nonsense, no matter what you say. You will lose to China.

In this case: The top labs voluntarily commit to embedded third party inspectors? Plot to ban open source. Regulatory capture. Also totalitarianism, presumably, and you’ll lose to China.

The only alternative would be to give open model creators money, and exempt them from all liability or responsibility for what they do, and sell our best chips to China. That’s freedom.

At some point, you need to stop taking these objections as mapping onto reality.

Joe Weisenthal: Any random restaurant regularly has third party inspectors checking, you know, refrigerator temperature or gas lines or raw meat handling. What’s the argument against having an AI “lab” equivalent? Alex Armlovich: Bank regulation is the best analog There’s limited public transparency: banks can’t be made to expose IP or publish everyone’s private banking data etc So the Fed & OCC embed permanent “resident examiners” in the most complex banks for continuous but confidential supervision roon (OpenAI): it’s totalitarianism Joe. William Isaac: Which Western-based OW model developer has the resources to support an embedded evaluators scheme? How would this not concentrate power even further? I don’t believe this is a sincere take. Dean W. Ball: yeah, you’re right, how could Nvidia and Meta ever afford this

Yes, such competitors have the resources to train frontier models but not the resources to allow for a few embedded evaluators, so voluntarily having embedded evaluators to check your power is instead a plot to concentrate power.

That does not mean that these concerns are never real. Some actions would indeed cause some combination of these four things to become more likely, or move us in such directions. And yes, we should take that into account.

But please, stop entertaining these same people saying this same thing every damn time anyone tries to do anything helpful. It is Obvious Nonsense, and Content-Free.

That also means not trying to placate or calm such folks, or address their concerns. You can’t. Their concern is that you want to take costly action to not die and perhaps not give them the maximally cool toys maximally fast. They are against this, and they are against it the same fixed amount no matter what.

I would not even dignify it by calling it a ‘conspiracy theory.’ There is no theory.

If your response is ‘such folks are reacting to Dario’s second and third proposals, not his first one’ then flat out no, that is not what is happening. You are incorrect.

You do want to address the underlying real concerns, to the extent they are legitimate. But do not make the mistake of trying to convince them you are doing this, or taking their reactions into account. They are a rock with these lines written on it. Period.

Yes We’re The Baddies How Did You Know?

The usual suspects from the previous section do have one redeeming feature.

They wear metaphorical skulls on their metaphorical uniforms.

If you must take the role of a cartoon villain, this is very good form.

As in:

No, seriously, this did not get deleted and he doubled down that he is not kidding:

martin_casado: I do wonder if dropping this a day after 9/11 was pre-meditated. I suspect so. Oh. Mostly just musing. The folks behind this are thoughtful if anything. And they have a penchant for the dramatic, massive unemployment, species extinction, etc. So dropping the messaging when we’re primed to be thinking about catastrophic happenings would be pretty inline with my experience with these folks. Yeah maybe it’s dumb. Maybe not. But it did cross my mind. I mean, if you’re comfortable using fear tactics around species extinction, you certainly be comfortable with this.

Sometimes People On the Internet Just Lie

I am done pretending ‘they don’t know the facts.’ They lie. And not well.

@jason (All-In Podcast, 870k+ views, QTing Dario’s call to Pace the Frontier): All of this because @openai decided they would EXPLICITLY ask thousands of software instances to find security bugs on the open web — while cos-playing a humans Best Supporting actor goes to @dwarkesh_sp for his soap opera recap of OpenAI’s hacking speed run. Here’s another concept: don’t create and scale thousands of viruses and worms and be SHOCKED when they actually cause damage! What a fucking farce this is.

There are plenty of understandable confusions about the HuggingFace attack. Accusing OpenAI of instigating the attack explicitly and on purpose? Yeah, no. That is not a misunderstanding Jason could possibly have reached, other than on purpose.

Dwarkesh Patel dutifully smacked Jason down while acting as if Jason is misinformed.

In case you were wondering if they are principled libertarians who don’t want the government on their side, well, no. As an example, Jason suggested responding to ‘the new models might kill everyone’ with requiring models be open sourced after a year. His solution is ‘we get to take your stuff.’

David Sacks Groks The Situation

Here’s a fun interaction, and no, as usual, you did not need Pangram to notice that the post was in the style of an AI, most likely Grok.

Charlie Bullock (showing David Sacks’s post ranting about Dario coming up as 100% AI in Pangram): IMO, this is still a major faux pas and you should not do it. David Sacks: Now I know these AI detectors are bogus. I run all my posts through Grok to fact-check them and suggest line edits but tell it not to change my writing style. That’s why my posts have edge. AI can’t cook like I do.

Okie dokie, sir. We learned something new about you today.

I strongly believe that it is indeed a major faux pas to post AI text as your own. It is fine to post it if it is clear that the writing is AI.

David Sacks Says Go Ahead

On substance, Sacks’s message boils down to: You have the advantage right now, and you’re the ones who say you’re doing something so dangerous, so you two (Anthropic and OpenAI) should slow down first, without worrying about antitrust laws and commercial consequences, and let the rest of us catch up to you. It would be good business, and then it would buy you goodwill.

David Sacks: ​The easiest way not to build superintelligence is for you to agree not to build it.

And yeah, okay, Sacks is overstating in many places including that one, but at the center of it is a pretty good point. Not that Sacks or his ilk would ever listen, they will read such a concession as weakness and foolishness and propaganda, and attack even more, maybe even call for antitrust action against them.

Indeed, Sacks himself says the companies should ‘agree’ to do this as a duopoly, which is a big no no without a waiver, and exactly the thing we have to dance around until the government does (less than) the absolute minimum and gives us the freaking waiver. What happened instead was that Anthropic did it on its own, and OpenAI followed.

You cannot simultaneously say ‘go ahead and do [X] in the illegal way’ and also say ‘without us giving you a waiver that permits [X].’ Well, I mean you can, Sacks did it, but it is not a good faith move.

That doesn’t matter. It is not about the mustache-twirling villains. It is not about fair.

It is about the fact that if you are about to do something existentially risky, maybe the first thing you should do is Stop It.

Lies and Confusions About Who Previously Claimed What

This keeps happening, and it will continue to happen, as part of the campaign of association to call anyone who warns about any downsides of tech as ‘doomers’ and then to associate them all with each other, to then dismiss all concerns.

Rachel (wrong or worse): Everyone who was telling me that climate change was going to end the world is now telling me that AI is going to end the world.

I do not know of anyone prominently warning about AI existential risk who previously predicted extinction or other unrealized major downsides from climate change. Not one case. Those worried about AI acknowledge that climate change is a real problem and try to be helpful in finding real solutions, without catastrophizing.

Christopher Ingraham (even more wrong): I think it’s more that the people telling me AI is going to end the world are the same ones who were telling me that NFTs would reshape the global economy

This is the one that blows my mind. There is almost no overlap between NFT evangelists and those worried about AI killing everyone. It is quite the opposite, as Yglesias and Gross say below.

Matthew Yglesias: These kinds of jokes are fun to make but I wish people would check before making them — the climate movement spent years *hating* on x-risk people because they felt it was a distraction from climate. The big NFT boosters are in the Andreesen “let er rip” camp on AI. Ben Gross: Fundamental confusion of two different tech communities: the people telling you that NFTs would reshape the global economy are the ones who think AI frontier development should be accelerated. The people telling you AI is going to end the world were also saying it five years ago. Jack: ah yes, the notorious group “everyone who talks about things I don’t know about and don’t like and who therefore must all be on the same team.”

The counterargument is ‘well they vibe the same to me, Jack’:

Arthur Baker: Out-group homogeneity bias! These duos each probably do read as culturally similar if you’re culturally distant enough.

There is actually a huge difference, but yes, a lot of people simply file everyone under ‘outgroup’ and then link them accordingly.

An especially fun one is to claim that all of this concern over existential risk is ‘new’ or that it needs to be ‘explained’ by some cynical reason, or even that it (lol) shows that they aren’t making progress on capabilities, are you kidding me.

Sabine Hossenfelder: fwiw I think the actual reason leaders of AI frontier labs suddenly seem to agree to pace AI development is (a) safety measures are currently so crap that they’re likely to end up in court if not jail, and they all agree that no one wants that
(b) they see no major new model advancement coming up soon anyway and need an excuse
(c) shift in public opinion the talk of existential threat is mostly there to keep you distracted and the stock market happy. Noah Smith: This is insane, they’ve been saying the same stuff for years, you just weren’t hearing about it or paying attention to it Derek Thompson: I feel like I’m taking crazy pills with how many smart people I read and respect are saying stuff like this. The idea that the AI labs are suddenly, only now, just in September of 2026, advocating for AI safety, and that they just came around to this position bc of sudden financial precarity, is just completely, utterly, conclusively, 100% wrong. Here’s Dario Amodei telling Ross Douthat in February that he agrees with the case for slowing down; that he’s in favor of “collaborating” internationally to organize a slowdown; that “I would be all for” a global slowdown. This was 7 months ago, when Anthropic’s annualized recurring revenue was rising faster than any company in modern history. He’s been saying stuff like this for years, when Anthropic was worth millions of dollars and when Anthropic was projected to IPO for trillions. it’s even stranger for her to tweet part (b) about OAI. do you honestly think that Sam Altman believes that model progress has stalled? He thinks Astra is a step sideways? he thinks solving Millennium problems is stagnation?! we’re just making stuff up here!

House Speaker Mike Johnson Wants To Lock Everyone In a Room

The reporting on this has often been absolutely terrible. Yes, Mike Johnson has correctly pointed out that the AI companies have an obligation to make their products safe, the same as every other product. In no way has Mike Johnson said the safety of AI is not also the responsibility of Congress or the White House.

Cheyanne M. Daniels (POLITICO, misrepresenting Mike Johnson): House Speaker Mike Johnson on Sunday said he would support talks with tech leaders on artificial intelligence, but insisted that developers — not Congress — are responsible for ensuring their products are safe. In an interview with CNN’s “State of the Union,” the Louisiana Republican argued Congress has already instituted AI guardrails and that “there is an obvious corporate responsibility that the people who are creating these models have to ensure that their products are safe.”

Congress has not already instituted meaningful AI guardrails. It is bizarre to be using that as a talking point. But yes, there is an obvious corporate responsibility to ensure that your products are safe, even in the sense of mundane safety. Certainly you also are obligated to make sure your product does not (checks notes) kill everyone on Earth.

At no point did he then say ‘and this is not my problem.’

One core objection of his is that while all the AI leaders say we need guardrails, different leaders all want different guardrails. This is a reasonable objection that can be fixed.

He actually said:

Mike Johnson (Speaker of the House-R): I’ve talked to the president about this as well. They [the AI companies] probably should be summoned together at the White House. And I think we need to go in a big room, close the door and sort this out. And I’ll be the one pushing for it.

Alas, Johnson falls back on Lose to China, but he does so in a newly balanced way.

Mike Johnson (Speaker of the House-R): We cannot put a moratorium on this because China will overlap us, and that’s the challenge. It’s national security balanced with the immediate security of making sure the models are safe. … We need to handle this new technology like we have others in the past and make sure we’re doing everything we can responsibly to also not smother American innovation. We have to do both things simultaneously. … We have to put some guardrails, some safety measures in place to ensure that AI doesn’t run away, that you know, we don’t have an out of control AGI. … We’re going to balance this. We’re going to have steady hands at the wheel as we do on every issue.

Words like ‘pause’ and ‘moratorium’ are scary to those with this mindset. I get it. And yes, if you waited long enough and China proceeded as per normal then China would potentially match and then pass us, although it would take longer than people think because the Chinese are mostly fast following.

Mike Johnson has no intention of having the House of Representatives in session debating things and passing laws at this time, including on AI, due to midterm elections. That is different from saying Congress has nothing to do with any of this. His reluctance to ‘rush in and pass a piece of legislation’ has multiple causes.

Cheyanne M. Daniels (POLITICO): Utah Republican Gov. Spencer Cox also appeared disappointed with Johnson’s response. Speaking on “Face the Nation,” Cox said Congress should be playing a “much bigger role” than it currently is. Spencer Cox (Governor of Utah-R): We’ve seen the reports. We know what’s possible. It’s very clear — and we’ve known this now for a little while — that the possibility of an existential event happening is growing and growing much more rapidly than anyone projected. And it’s really important that the government gets involved.

Well said.

Here is Mike Rogers, the Republican running for Senate in Michigan:

Mike Rogers: My experience as Intel Chairman taught me the time to stop a crisis is yesterday. The first role of our government is to keep us safe. Following Dario Amodei’s warning, I believe we must slow down development of AI in order to allow the government and its partners to catch up and put stricter protocols and guardrails in place to protect us from the rapid advancements we’ve seen. In the Senate, I will make it a top priority to pass the national security framework we need on AI.

Here is the Minority Leader:

Hakeem Jeffries (House Minority Leader-D): [Our caucus will meet on Tuesday] and a high priority will be action on artificial intelligence. ​ We should take decisive action now so that we can slow down, as the CEOs have recently acknowledged, slow down the pace of development in order to protect the American people and ensure that A.I. is proceeding safely.

This is exactly the right priority. You want to ensure the ability to intervene, including the ability to slow down.

Donald Trump is Not Tired of Winning

Donald Trump has already ‘paced the frontier’ somewhat, slowing down releases of Fable and Sol, and even yanking Fable from the market. Donald Trump’s administration was willing to act when the threat was concrete and in front of them.

He is willing to put guardrails on AI. But he’s all about keeping the vibes good, except where he’s all about keeping the vibes bad. Gotta have the right vibes. And Trump’s not yet buying this ‘existential risk’ thing, that’s ‘negative forces.’

Donald Trump (President of the United States): Well I say this. ​We’re leading China on AI. We’re the most sophisticated country in the world and frankly I want to keep it that way, because whoever wins AI wins. And we can put guardrails, and we can do this and that, but I think you have a lot of negative forces that are bringing it up, that shouldn’t be bringing it up, and they’re bringing up things that won’t happen.

How dare ‘negative forces’ ‘bring it up’ when it ‘won’t happen.’ Why won’t it happen? It won’t happen. You’re bad for the vibes, dude. Trump is willing to do what it takes, so long as you play nice and have the right vibes, which can in some contexts mean alarming or bad vibes, big scary vibes, the worst vibes, but not quite here and now.

The misleading headline chosen for this was ‘Donald Trump rejects calls from tech bosses for an AI slowdown’ and that he ‘denounces demands for regulation’ but that is terrible reporting. It is not what he said. I would caution against reading too much into Trump’s statement. You can find the full clip here and I encourage you to watch it from about 7:30 when the question gets asked.

roon (OpenAI): this is rather bad reporting. nowhere in the transcript does he “reject the slowdown”. he says something about guardrails being okay and the need to win and then starts rambling. sounds pretty much on board to me

Trump is asked ‘should AI be regulated or slow down?’ From his perspective those are trigger words that mean something very different than ‘embedded evaluators’ and ‘not build Dyson Spheres in 2029,’ also he is unlikely to have been properly briefed. After the above passage, including agreeing that ‘we can put guardrails,’ Trump pivots to ‘look at all the factories we are building.’ What Trump is trying to do here is head off the data center protestors and full pause advocates and negative vibes. Then he talks about which reporters look better.

Trump then reiterated this on Truth Social, which is easy to misread but you have to actually read the words:

Donald Trump (President of the United States): ​The only control or “guardrails” that AI needs is a STRONG AND SMART (High IQ!) PRESIDENT, and the U.S.A. has that, in spades! The Trump Administration has stopped AI “people” from doing bad, or potentially bad, “things,” like Dario (Anthropic!), who is now pretending to be a “perfect little angel” – and we will continue to do so! We already have tremendous CRIMINAL and REGULATORY power over these companies! There is a SICK conspiracy going on against AI and Data Centers, and the only one that is happy about it is China. WHOEVER WINS AI, WINS! We are leading China, and all others, and will continue to do so. Conspiracy Theorists, Treasonists, Traitors, and Leakers, BEWARE! Thank you for your attention to this matter!

Trump is mad about the anti-data center thing and the potential conflation of the two, and about the implication that things might go wrong on his watch, and taking a potshot at Dario Amodei for once again bringing the bad vibes. But obviously you should not interpret this as saying ‘the guardrail is literally that I am the President and therefore we do not need to, what’s the word, actually ever do anything.’ The guardrail is that he will choose the guardrails, and so on. Keep those vibes good, folks.

This is also standard muscle flexing, which may or may not have anything to do with threatening antitrust actions. I doubt it, but such strategies thrive on strategic ambiguity. It’s hilarious how much Fable takes these ‘criminal threats’ literally.

Trump is indeed strongly opposed to anything like a full pause, and is a big fan of the data centers, but none of that is news. It also could change, and fast.

Thus, after a brief period of hopefulness, the chance of a federal AI safety bill this year is back down to 18%, which I interpret as correctly saying ‘not without another incident, and maybe not even then.’

We Must Avoid Polarization on AI at (Almost) All Costs

I endorse this.

Eliezer Yudkowsky: There are currently zero things more important than “don’t die to AI” becoming a bipartisan project rather than a Democrat-polarized issue. If you have any political capital you can spend on this, do it now, I beg of you. Life or death. If you are a Democrat: I beg you, do not throw out barbs at Republicans about this issue. Ask them to join with you on this common interest. If you are a Republican: Defend your country, your people, your home and your family from death.

How Will We Know If They Actually Paced?

I disagree with Daniel in that I am very confident this is not ‘regulatory capture.’

I agree with him that we should be very concerned that it is kabuki, or that the labs do not meaningfully slow down, or they might only slow down in the places where prosaic safety would have stopped them anyway for prosaic reasons.

We don’t know what the counterfactual pace of progress is, and it is hard to know what progress is supposed to look like when cashed out into practical and real-world capabilities and applications.

Daniel Kokotajlo: We’ll be able to tell if this is regulatory capture or not by whether the pace of progress at the frontier actually slows down (which would allow other companies to catch up). If a year from now we look at the trendlines in A\ and OpenAI capabilities progress towards recursive self-improvement (e.g. ECI, time horizons, coding uplift), and there’s no visible decrease in slope, then probably we were cheated. (As I hope is clear, I am somewhat concerned that that this will happen. I’m speaking up now in the hopes of shaping the developing situation to be *actual* pacing of the frontier instead of safety washing regulatory capture.) Drake Thomas (Anthropic): Hm, it seems very plausible to me that the unpaced default is a rapid increase in slope from RSI and “keeping ECI slope the same” is the result of aggressive pacing. I would feel a lot safer if I knew that we would just maintain current trends. Daniel Kokotajlo: I also would feel a lot safer in that case, but nevertheless I’m asking for more. I wanna see my trendlines bend downwards. It’s too hard to verify corporate claims otherwise — e.g. if the trendlines have continued just as fast as before, and Dario says he could have killed me by now but didn’t because of all the safety work and pacing they did, how can I tell whether he’s right vs. it’s all safety-washing and things are still going max speed? Might be confusing and hard to tell.

This can also be viewed as a question of what pacing means. Does it mean ‘this is costly but worthwhile’ or does it mean ‘this is necessary and we would have to do it anyway, and thus it is worthwhile’?

It can be both, or that can become a real disagreement. What is the counterfactual? If you couldn’t proceed because your product was too misaligned to use internally, did you actually ‘pace’? Kinda yes, kinda no, right?

The embedded evaluators are designed to be able to address this question, by reporting on what steps have been taken and what the counterfactual might have been. This has to become a ‘trust but verify’ situation.

If we can keep things on the current trend lines, certainly that is way better than breaking the trendlines upwards into imminent recursive self-improvement. We would go from remarkably little chance to at least some chance. But I agree with Daniel that if we see ECI staying on-trend, and that is an accurate representation, we are still moving way too fast. The last year was too fast, and we are paying even the prosaic bills now in terms of the crisis of misalignment.

The Real Frontier Is Internal Models At Top Labs

Roon is correct that misuse risks are a problem but internal deployments going haywire are now the primary concern. We are launching swarms of 10,000 Astra-variant agents at plausibly impossible tasks like Millennium problems a week after training starts. It’s not a great leap to imagine how that might go horribly wrong.

roon (OpenAI): Pacing the frontier applies first and foremost to internal model development. The risks from raising the capabilities waterline in internal deployments are much higher than misuse risks. a new model can outsmart you in weird ways, exfiltrate its weights etc. *training* becomes costlier, and deployment will stay the same: you should expect the gap to shrink

This is what must be paced. This is where the battle will be won or lost.

We can finally get to work, and have our chance. Don’t waste it.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论