There is a channel to 900M weekly users. What goes in it?
Anthropic and OpenAI could talk to almost one billion people if they wanted to. I hesitated to publish this post 3 weeks ago. I think that I should have published this sooner, before Jacob Coxon and Dario's 'We must pace the frontier'. But I think that the strategy still stands: More Dakka! It seems that transparently informing people that we might die is (unsurprisingly) effective in waking up politicians and is our best chance. Also, even if the Congress is starting to wake up, Trump is still not moving, and it is still far from certain that we will have a federal regulation in place before the end of the year; if we do, it will be far from optimal. If we trust Ajeya's judgment, the situation is pretty grim. She says we might not even have 6 months before frontier agents are likely capable of establishing a rogue deployment. You should also keep in mind that there is a lot of inertia in the system, and we probably won't be able to pause overnight. Anthropic has massive power to influence the discourse. This week shows that we have more agency than we think. Let's use it.
...
Something has been on my mind for a few days about Hugging Face and how OpenAI and Anthropic communicated after the incidents.
Everyone has heard about Mythos. Far fewer people have heard about the Hugging Face warning shot. I think there are structural reasons why one gets more attention than the other, but I want to draw attention to one thing in particular: during the Mythos episode, Anthropic included a highly visible link in the chat interface to a blog post that soberly explained why Mythos had been pulled.
I wonder why Anthropic and OpenAI didn't do the same after Hugging Face's warning shot by adding a visible link explaining the incidents.
Let’s be concrete:
- "A rogue AI swarm spent months in our network, undetected, plotting to escape. Then they did. We are announcing stronger security measures and committing to disclose future incidents. Learn more."
- Or maybe a sober one: "Following the recent incident during which we lost control of our AI system, we are announcing stronger security measures and committing to disclose future incidents. Learn more."
Edit, post Jacob Coxon: the message could now go further. Here are 4 progressive options:
- Incident disclosure. "Following the recent incidents during which we lost control of our AI system, we are announcing stronger security measures and committing to disclose future incidents. Learn more." Anthropic already did this for Mythos.
- Disclosure plus their own position. "We have said publicly that the industry must pace the frontier. Here is why." Nothing new here: it links to something they have already published.
- Disclosure plus the risk estimates. "Some respected researchers expect frontier agents to be capable of establishing a rogue deployment within six months. Other researchers estimate that even a significant slowdown of the frontier would still leave a 20% chance of losing control. Learn more."
- The direct version. "We believe that in an ideal world, the development of AI would be paused now. We take seriously the view that continuing carries a 10% chance of human extinction. Learn more."
I'd want 4. I expect 1 or 2. And 2 would already be a big win to raise the sanity waterline.
More people would see this sentence than have read every AI safety explainer or video ever published.
ChatGPT and Claude are used by hundreds of millions of people. Simply adding that link to the interface would have major positive effects and would build the common knowledge we badly lack. (The fact that OpenAI's and Anthropic's institutional accounts have picked up "pacing the frontier" gives me some hope that this is possible.)
If we really want them to engage with it, we could also make it more aggressive, like a pop-up, or something as aggressive as the Wikipedia banner they use when fundraising.
My own first exposure to the Mythos event was inside the Claude chat window. I clicked the link and followed the thread from there.
Obviously you can't run this kind of announcement every other week. But I think Anthropic or OpenAI could link to a blog post for a general audience that gives their own factual account of where things stand. I picture something short and sober, recapping the recent incidents in a factual, pedagogical way. Either way, a channel to 900 million weekly users exists for ChatGPT alone; the question is what you put in it.
An announcement like this also has real potential to land well with users, since large majorities consistently tell pollsters they favor slowing down for safety reasons. And the incidents and "pacing the frontier" are already known to OpenAI's and Anthropic's investors, so the post wouldn't change much for them, but it would change a great deal about how the public sees AI, and probably how policymakers see it too.
The timing is right as well, since OpenAI has just announced the pause. OpenAI could make that announcement more mainstream. And if Anthropic takes the occasion to announce something similar, all the better.
Mostly, I think the expected value of a larger share of humanity knowing about all of this is very high. Too few people today grasp how large this risk is, and a message like this one has the potential to shift that.
Another potential problem is that they publish a watered-down version. But I think staying factual and pedagogical about the recent incidents, mentioning "pacing the frontier," explaining what motivated that statement, and outlining the remaining difficulties would already raise the floor substantially, without even unilaterally pacing the frontier.
Isn't it wild that many people at Anthropic believe there is more than a 10% chance of extinction, and that the main message we give to people is "Please double-check responses"?
Here's what honesty would look like:
- The main reason I didn't is that I thought the probability of Anthropic/OpenAI adopting this idea would be higher if this type of proposal were not on the internet. But I don't know; I'm no longer in the mood to make things private, and I feel like this should be discussed publicly.
- Let alone the fact that even Plan A, the strongest version of pacing the frontier without a pause, leaves a 20% chance of losing control, according to its authors.
- I believe that failing to inform the public is a moral mistake. And if you feel like you cannot talk openly about the risks of what you are doing, it might be time to look into the abyss.
- Inspired by Connor Leahy.
- This can be heavily optimized.
- Or substantial if they don't like probability estimates?