A ‘Morally Binding’ White House Accord on AI Safety
The leaders in AI were invited to the White House. We left with a White House agreement that is nonzero Actual Progress rather than a step backwards.
The key to success, in many situations, is to call the whole operation something else.
When you have one side that cares mostly about vibes, and the other that cares about the substance, this suggests a deal that can be struck.
Suddenly everyone agrees on everything. Works for me.
That doesn’t mean peace in our time. The next fight is already ramping up, as we see signs that they will make another attempt at an insane, maximally bad moratorium during the lame duck session.
Table of Contents
- Look Who’s Coming To Dinner.
- Let’s Do Lunch.
- I Think It’s Morally Binding, Yeah.
- Everyone Who is Anyone.
- The White House Accord on [Artificial] Intelligence.
- The FTC Investigates.
- We’re Going To Need a Stronger Regulatory Regime.
- [Artificial Intelligence].
- Money, Dear Boy.
- They Are Going To Try This Moratorium Insanity Again During the Lame Duck Session.
- The Quest for Embedded Evaluators.
- Hugging the Face.
- Reinforcement Learning from Heartland Feedback (RLHF).
Look Who’s Coming To Dinner
After the summit between America and China, Trump invited Dario Amodei to a 1-on-1 private dinner. Trump said he would talk about the whole ‘slowing down’ thing and advocate for ‘let’s go’ and ‘let’s win.’
Dario Amodei got the training he needed to take on that challenge. This was a key moment to cut through the haze, as there were a number of important things here that Trump clearly did not understand. It was also, perhaps more importantly, a chance to mend the relationship on a personal level.
Dario also met with Thune for 20-25 minutes.
It looks like it worked. Dario did at least some things right, provided he did not give anything important away that we are not tracking:
Ben Brody: “Dario has been fantastic!” Trump says of Amodei, with whom the admin has had a very troubled relationship. Amodei downplayed any distance between him and Trump. JgaltTweets: Trump: “Dario’s been, actually, great… Dario has been fantastic, and we had dinner the other night, and he agrees with everyone. I mean, we all agree.”
In the clip that quotes, from Rapid Response 47, you see a CNN reporter trying to ask Dario a question about whether this will be enough, and Trump cutting the reporter off saying ‘Hold it, hold it. CNN Fake News. She’s Fake News. The worst.’ Great move, since I am guessing Dario really didn’t want to have to answer that one in front of Donald Trump and they got to do a bit together.
This will probably flip at least five times in the next twelve months.
Let’s Do Lunch
It still wasn’t the best possible dinner, which we know because (I can’t believe I have to type this) we can see the seating charts and Trump is not close to Anthropic cofounder Tom Brown or CEO Dario Amodei, and indeed you’d have to pass through Brad Gerstner and Richard Walters:
Andrew Curran: For the many people asking where Dario is:
Whereas Trump has direct access to Musk, Huang and Zuckerberg, and Sacks is also rather central on the other side.
However, Brockman wasn’t that central either, and Anthropic got two slots, so maybe we should not read too much into such things beyond what we already knew.
Trump is very much still on the pro-AI warpath, even if he insists we not call it that, especially emphasizing the benefits of data centers.
The thing about the pro-AI warpath is it is fully compatible with taking basic safety measures and putting up guardrails. Having the whole thing blow up helps no one.
So you can, if desired, do the responsible thing, and frame it as the warpath.
I Think It’s Morally Binding, Yeah
Trump said there’s going to be ‘tremendous self-regulation,’ by which we mean jointly agreeing on regulation.
This is part of there now being a White House Accord on [Artificial] Intelligence.
Jake Sherman: [Speaker Johnson], on Squawk Box, says “a little oversight” transparency and external auditing on AI would be appropriate. Andrew Curran: Speaker Johnson said after the meeting that everyone signed The White House Accord on [Artificial] Intelligence, A Joint Commitment on [AI] Responsibilities. – voluntary commitments
– robust internal control
– layers of internal and external review ‘The White House and Congress will continue to guide the industry’s development in a safe manner’ … Reporter: ‘Mr President, this accord that everyone signed, is it binding in any way?’
President Trump: ‘I think it’s morally binding, yeah.’
Hahahahahahaha.
The headlines often actually say ‘morally binding’ this is too perfect.
In practice the provisions are simple enough, and basic enough, and obviously worthwhile, such that I expect this to be followed.
It’s still an important document. The real line that matters is easy to miss.
Everyone Who is Anyone
Yesterday I mentioned that the planned SAFA group to set AI standards, led by Google, OpenAI and Anthropic, was actively opposed by Meta, xAI and Nvidia.
This agreement is different. Meta, xAI and Nvidia all signed on the dotted line.
Microsoft and Amazon were present and did not visibly sign. It is unclear if they chose not to sign, or were not considered sufficiently worthy of inclusion.
The White House Accord on [Artificial] Intelligence
Here is the full text, with an external auditor for everyone, signatories include Sundar Pichai, Dario Amodei, Mark Zuckerberg, Elon Musk, Jensen Huang and Greg Brockman (since Sam Altman was at Dev Day):
White House Accord on [Artificial] Intelligence Joint Commitment on Frontier Responsibilities In order to build a positive future for the American people and the world, we believe every company is responsible for developing its own technology safely and in a way that builds trust with customers and the public. This starts with every company that is training and deploying frontier models having robust internal processes and controls to ensure that their technology behaves as intended and that any issues are promptly identified and resolved. Therefore, in addition to any other precautions, we believe each company should implement the following four layers of controls and audits:Together, these steps will give each company, its customers, and the public confidence that the technology is operating as intended. The participating companies will meet regularly to establish standards and best practices to improve the safety of their systems. Over time, it may make sense to codify these steps into laws or regulations. Regardless of whether this is required of companies, we believe that implementing these controls and audits is critical to ensuring a safe future for everyone, and each of our companies are committed to doing this.
- Implement robust internal controls to monitor the capabilities and alignment of its models during training and deployment around areas like cybersecurity, biosecurity, and chemical threats, and to ensure that its models do not hack or access technical systems in unintended ways.
- Empower an internal team to ensure all of the controls, monitoring, and detection are operating as intended, and that any issues are remediated.
- Partner with an independent external auditor or evaluator to carry out independent assessments of whether the controls, monitoring, and detection are operating as intended.
- Designate an independent committee of the board of directors to oversee and receive reports from the teams operating the controls and the internal and external auditors and evaluators, as well as to ensure any issues identified are remediated.
The other three things on the list are good things, but I mean can you imagine a company standing up and saying they weren’t doing those three things, given that such auditors and evaluators exist?
- “No, we don’t have robust internal controls monitoring the capabilities and alignment of our models during training and deployment around areas like cybersecurity, biosecurity and chemical threats, and to ensure that our models do not hack or access technical systems in unintended ways.”
- “No, we do not have an internal team empowered to ensure all of our controls, monitoring, and detection are operating as intended, and that any issues are remedied.”
- “No, we did not designate an independent committee of the board of directors to oversee and receive reports from the teams operating the controls and the internal and external auditors and evaluators, as well as to ensure any issues identified are remediated.
I can see the third one not being done, and using some other group, but okay, sure, it will be members of the board.
What matters is whether that committee actually cares, and pays attention, and gets informed, and acts. As referenced below under Hugging the Face, OpenAI had exactly this committee, but they either were kept out of the relevant loops or failed to act.
None of that is what matters here. Notice this line:
The participating companies will meet regularly to establish standards and best practices to improve the safety of their systems.
That is, in practice, close to an antitrust waiver.
I repeat. We at least kind of have an antitrust waiver, we just don’t call it that.
It’s not a formal waiver. In theory the DoJ or FTC or various others could still sue under the Sherman Act and other laws.
This is still a large de-risking. Yesterday I reiterated that I did not in practice think that antitrust was a major risk, and this was a ‘if he wanted to, he would’ situation. That is doubly true now. The White House could in theory go after anyone at any time for any reason on a whim, and has tons of sources of leverage, but in particular ‘go after the AI companies for having the discussions and setting the standards they agreed at the White House to meet for and set’ is not the thing to worry about.
You could make the case both that what was done here was already legal, and that it doesn’t mention the things like pacing that would be illegal. That strikes me as a far too literalist reading, and also discounts both that there was a lot of previous talk about how even discussing standards would be illegal, and the practical cover granted.
A lot of this is that the true goals are not in conflict. Trump does not want AI progress to stop, and panics when he hears ‘pause’ or ‘pace the frontier’ or ‘slow down,’ but when Altman and Amodei say ‘slow down’ they mean from ludicrous speed to only extremely fast.
Concretely, say we are on a bus called ‘the economy and AI progress and beating China’ or whatever that will blow up if we go below 50 miles an hour, but also we are currently hitting the gas so fast we are now going 100 miles an hour on pace for 200+. Both sides can get what they want at the same time.
The FTC Investigates
The FTC has issued a CID, confirming that there is an industry-wide probe that was opened this summer, with demands going out for information from at least OpenAI, Anthropic and METR, with at least some relation to the HuggingFace Incident.
We don’t know what they are investigating, but given the timing and targets one presumes that they are investigating hacking and the failure to make the models safe, rather than investigating recent attempts to fix the problem.
We’re Going To Need a Stronger Regulatory Regime
Ezra Klein talked to Bill Gates this week, with the headline description ‘Bill Gates thinks AI alarmism hasn’t gone far enough.’
Ezra Klein: Bill Gates thinks A.I. alarmism hasn’t gone far enough. He believes the years ahead will be marred by catastrophic cyberattacks, bio-terrorism and mass job loss unless the government acts, because he thinks the idea that the A.I. industry can just self-regulate is “insane.” I was struck in the conversation by how genuinely afraid Gates seemed to be that we are not ready for the consequences of what we are building, and how emphatic he was that so many of the people in positions of authority are in complete denial about what he thinks is about to happen.
Self-regulation as a plan is indeed rather insane. It still is better than none at all, and can serve as a first step.
Gates was asked, for example, whether ordinary laws and capitalist incentives would be able to handle this, and Gates responds correctly with ‘I can’t believe you asked me that,’ and then points out that not only is AI uniquely dangerous, this is simply not how society has ever dealt with actually dangerous things. We don’t say ‘well sure make your products unsafe, if you do then people will sue you.’ So, no, then.
[Artificial Intelligence]
The other supposed agreement was a distinct executive order to change the name of AI to [AI], which Trump claims the lab leaders signed off on. I’m sorry I have this tick, I try to type [artificial intelligence] and instead I end up with brackets.
Andrew Curran: President Trump ‘We’re going to be signing a document today at about five o’clock, renaming Artificial Intelligence, because it’s not artificial, we all agree on that, and we’re going to be renaming it [Artificial Intelligence]. Officially renaming it.’
I am very curious if Trump plans to now get mad every time Dario or Altman or Musk says the words ‘artificial intelligence.’
Trump is also claiming he is going to start talking about ‘Artificial News’ to describe the media, which is about the time I realized I kept typing the brackets. Shall we say.
So why would Trump go this hard on something like this?
The obvious explanation is this is Vintage Trump. He’s all about renaming things, and finding nicknames, and invoking vibes. He has a been a world class vibe invoker, nicknamer and term associator, especially in the 2016 election. If you think he chose the most annoying possible name due to namespace clashes, it’s probably because he was optimizing for that on some level. There is method to the madness.
You can also see it as a ‘bend the knee’ moment, the way various regimes and cultures often insist people affirm absurdist things. If you can get the makers of AI to start calling it [AI] instead, in a sense you own them. They have shown loyalty. And this is a way of weeding out those who won’t play along. Will Dario say AI or [AI]?
Money, Dear Boy
There is also an alternative explanation that is dumb and corrupt even for 2026, that has been put forward by Adam Cochran. As additional suggestive evidence, two of the three alternative options for the renaming that were in Trump’s initial poll, and both of the final two after he restarted it under false pretenses, started with S.
As counterpoints, the term ‘superintelligence’ was already being used as marketing by Meta and talked about by others, so these investments were smart anyway, and the timing would mean that this was a plan months in the making, which would not be Trump’s style, and would be in conflict with this looking strongly like a reaction to the HuggingFace incident and Pacing the Frontier.
Is it possible that this is part of the motivation for the renaming drive? Of course, yes, we absolutely live in a timeline roughly this dumb.
My strong presumption, however, is that primary causation runs the other way. Any insider trading is mostly a free action given plans that exist for other reasons, rather than a driving force causing the changes. I don’t like it, but I am not that mad about it.
Adam Cochran (adamscochran.eth): SCOOP: Trump’s “Super Intelligence” Scandal: I believe Trump’s “SI” Executive Order was ANOTHER criminal plot to enrich the Trump family. Insiders seem to have profited MILLIONS off of .si domain names before his Truth Social posts. It’s no surprise that .AI domain names are a hot commodity and almost all taken by squatters. But what wasn’t? .si domain names. .si is the domain extension of Slovenia (coincidentally where Melania and Barron Trump are both citizens of) On September 19th Trump made a random Truth Social post about renaming “Artificial Intelligence” to “Super Intelligence” without any clear reason. This set off a flurry of people buying and registering .si names related to AI. But it wasn’t the first time… The .si registry has around 55/day registrations in 2023-2025, but in 2026 the numbers started to pick up. On June 24th they saw more than 400 in a single day. This volume spiked with more than ***7000*** new domains registered in July. And a sudden flurry of buying .si domains in the aftermarket. AI related .si names started going for tens of thousands of dollars. … Since Trump’s announcement .si domain have seen *MILLIONS* of dollars in turn over. Much of it going to domains that were registered in the last 60 days before the announcement. TL;DR:
-Trump made the SI executive order to try and force companies to brand as “SI” instead of “AI”
-In the past 60 days insiders bought thousands of .si domain names
-They’ve profited hundreds of millions of dollars so far.
They Are Going To Try This Moratorium Insanity Again During the Lame Duck Session
Don’t let them. They are, for reasons I actually cannot fathom, going after exactly the worst possible preemption law, which they may try to sell by combining it with the promised codifying of existing industry practices mentioned in yesterday’s agreement.
Ben Brody: Breaking: The focus of preemption in the Thune/Cruz/Klobuchar AI talks? State laws on catastrophic risk like nuclear concerns. Way narrower than Cruz’s decade-long moratorium of last year but could still cause worries for Dems on handling of CA etc.
Preemption on catastrophic and existential risks, and ensuring the safety of frontier models, is a very bad idea.
In theory preemption could be the correct move if it was part of a sufficiently robust and trustworthy Federal response on such matters, with expectation of real enforcement of both letter and spirit of the laws involved, and that we would reasonably respond to changing developments.
In practice, none of that is a reasonable expectation for something that might emerge from a lame duck session. We would be trading in all the state laws, and all potential state laws, for something the White House would likely only selectively enforce in an ad hoc manner, that was largely crafted by Ted Cruz, and which you would be unable to amend later because they already would have preemption in hand.
I am not saying there is no way to ‘get to yes’ on such a deal in theory, but there is no way they are going to offer anything reasonable in practice. Just say no.
Whereas, when it comes to things like child safety, or disinformation and deepfakes, or other mundane and ordinary consumer harms, a unified code makes perfect sense and the patchwork of laws could turn into a Kafkaesque nightmare and will often be about protecting rent seeking and preventing diffusion, so preemption would make perfect sense alongside a sensible set of rules.
So, of course, we get a narrow provision, allowing states to do all the disruptive and negative things, but stopping them from doing the things that might make us not die.
Don’t let them do it.
This will be the key test of the OpenAI supposed Heel Face Turn.
If OpenAI and its lobbying arms support bills with such preemption provisions, and some signs suggest this may be their plan, then they will have betrayed us and played us for fools. It would put a lie to their entire Heel Face Turn.
Be prepared to update accordingly. And let this be a warning. Don’t try it.
The Quest for Embedded Evaluators
Apollo is ready to throw its hat into the ring, as a ‘METR-style’ evaluator. Their new post argues that embedded evaluators are feasible, already proven effective, and they are necessary to have a real shot at catching events like the HuggingFace Incident.
Apollo Research: New Post: External testing needs embedded evaluators with employee-equivalent access. Recent incidents mostly occurred during model development and internal evaluations, while current third-party evaluations happen before public release. Embedded evaluations can close that gap. The industry should be working toward employee-equivalent access for embedded evaluators. The key questions are often about process: when a monitor flags something, who reviews it and who can stop the run? Checking that requires access to people. Marius Hobbhahn (CEO Apollo Research): Apollo has worked over the last couple of months on shifting our evals program to be an embedded evaluator (before Dario post and HF incident). Without employee-equivalent access it is hard to make meaningful safety assessments. We’ll publish more details in the coming weeks!
Hugging the Face
In case you needed extra motivation for why we need those Embedded Evaluators and auditors: To the great surprise of absolutely no one, months earlier, and got ignored because the show had to go on and the models get out on time.
This is confirmation but I’m not sure I’d call it ‘news.’
Sheera Frenkel, Dustin Volz and Dylan Freedman (NYTimes): In emails, the employees said they worried that OpenAI’s newest artificial intelligence models were not being appropriately monitored during testing to gauge the technology’s sophistication and to secure the models, according to messages viewed by The New York Times. In response, OpenAI executives told the employees that the tests needed to move forward as quickly as possible to release the A.I. models on time. No additional security protocols were instituted, said the workers, who were not authorized to speak publicly on sensitive matters. … The exchanges between OpenAI employees and executives — which have not been previously reported — were part of a pattern where the San Francisco company did not prioritize security, according to employees and independent security researchers. Nathan Calvin: It’s remarkable that the employees felt strongly enough about what happened here that they were willing to take the personal and legal risks associated with talking the NYT. Either the safety and security committee of the nonprofit board was not notified about these warnings, or they were and didn’t intervene. Either option seems very bad. This reporting also seems extremely relevant to the question of whether OpenAI is fulfilling the legal promises it made to the CA and DE AGs during its nonprofit restructuring to put safety and security before commercial interests.
It is amusing to see this in The New York Times, I mean yes, true enough I guess (Daniel Kokotajlo left a while ago and is not one of the employees this time around)?
If anything, I find this part more troubling:
Sheera Frenkel, Dustin Volz and Dylan Freedman (NYTimes): In July, researchers at the security company Hacktron said they told OpenAI about how they had found a way to break into the company’s systems with the help of an A.I. model created by its rival Anthropic. OpenAI initially dismissed their findings, they said. In a shared channel on the messaging platform Slack, Mr. Stuckey of OpenAI wrote that it was “pretty sad” that Hacktron’s researchers had gone to such lengths to demonstrate the company’s vulnerabilities, according to copies of the communications seen by The Times. “We just felt like they were angry at us,” Mohan Pedhapati, a Hacktron researcher, said of OpenAI. … Mr. Stuckey later apologized to Hacktron and OpenAI awarded the researchers $6,500 for disclosing the flaw.
If you are a white hat hacker, and you break into OpenAI, and are careful to not actually access any confidential information and instead turn in the vulnerability for a $6,500 bug bounty, they should be thanking you profusely and thinking you were severely underpaid. If you come away thinking they are angry, things are not good.
There is another vulnerability described later, where OpenAI paid out $500.
Sheera Frenkel, Dustin Volz and Dylan Freedman (NYTimes): OpenAI gave $500 to the group for its work, which Mr. Wardle said was low compared with what he would expect from other companies given the severity of the flaw.
I do not understand why these figures are not more like $50,000, or perhaps $500,000.
Reinforcement Learning from Heartland Feedback (RLHF)
Trump promises to name an AI czar in the next few days, after which it sounds like he intends to go off on the campaign trail in the traditional ‘we don’t need a government right before an election’ period.
That means rallies. I am very, very curious what happens when Trump’s ‘AI is great, in fact it’s SUPER INTELLIGENCE, and data centers are awesome and will save you’ message interacts with the American people. What will Trump do when he hears the reactions of the crowds? One of his superpowers is noticing when the crowds are not vibing with what he is selling, and pivoting to selling something else.