AI #183: Pre Post Mortem

Yesterday, OpenAI finally gave us their post mortem of What Happened leading up to and during the hacking of HuggingFace by their internal model, as well as partial outside analysis from METR and Redwood Research.

The reports are a doozy. I am only beginning to work my way through them. I would have pushed the weekly to cover that today, but I need more time, so I plan to start coverage of the post-mortem tomorrow, along with related other events.

I’ve also spun out a few other discussions, including on ‘aligned to whom,’ on cooperative alignment things and on when you can trust lab messaging, as part of the new direction of more focused posts on AI topics that I polish a bit more.

Table of Contents

  1. Language Models Offer Mundane Utility. Check your facts.
  2. Language Models Don’t Offer Mundane Utility. How much would you pay?
  3. Huh, Upgrades. ChatGPT can access your iMessages.
  4. Get My Agent On The Line. Also get some sleep. You can’t go on like this.
  5. Deepfaketown and Botpocalypse Soon. What makes AI content repulsive?
  6. Cyber Lack of Security. Chinese hackers broke into the Federal Reserve?
  7. Reinventing OpenAI. Alex Heath covers OpenAI’s response to HuggingFace.
  8. They Took Our Jobs. Bill Gates warns of ‘economic catastrophe’ and more.
  9. What Is The Law. Bottom tasks automate and fall out, and winners take most.
  10. Job Retraining Programs Don’t Work. Never have, probably never will.
  11. Get Involved. OpenAI Foundation and SecureBio are hiring.
  12. In Other AI News. Fable is getting criminally underused, question is why.
  13. Show Me the Money. Anthropic supervoting shares, Nvidia buying HuggingFace.
  14. Quiet Speculations. When will then be now? Soon. Problem for future Earth.
  15. If You’re Not Going To Take This Seriously. No one credible can do the modeling.
  16. Quickly, There’s No Time. Peter Wildeford brings us his timeline updates.
  17. The Quest for Sane Regulations. Mixed signals.
  18. Don’t Panic. A brief look at the history of American moral panics.
  19. Pacing the Frontier. If we wanted to do it, how would we do it?
  20. Chip City. Effective data center arguments, new chips IN SPACE? Really?
  21. The Week in Audio. Bostrom, Altman, Kokotajlo.
  22. People Just Say Things.
  23. Rhetorical Innovation. A look back, some looks forward.
  24. Mundane Incremental Alignment Is Worthwhile. Necessary, but insufficient.
  25. New Blog, Who Dis. Dean Ball launches a blog within OpenAI.
  26. Other People Are Not As Worried About AI Killing Everyone. Factorio.
  27. The Lighter Side. All I remember are the Titans.

Language Models Offer Mundane Utility

I agree with Owain Evans that AI is now excellent at mundane fact-checking and related styles of research, greatly outperforming pre-AI humans. The surveillance concerns are real, but mostly this is great in practice, and I worry many are missing out because of the 2022-era hallucination rates.

Be Delta Airlines and set different ticket prices for every passenger in real time, or be Uber and quote different customers $76 and $24 at the same time for the same trip. You really do have to be careful about sending signals that cause airlines or hotels to jack up the price on a specific trip on you in particular. There are solutions.

Language Models Don’t Offer Mundane Utility

How much would you need to be paid to give up AI, or various other things?

These numbers are remarkably small, especially the medians, and especially online search, especially if you also couldn’t ask others to do it or use generative AI. One could argue there is some value in getting a refresher or cleanse, as an experience, but the average value seems far higher than these numbers suggest.

This calculation leads to:

Needless to say, even if I wasn’t trying to keep up with or write about AI, you’d have to pay me quite a lot to give it up for a month.

Google needs to get its act together, for so many different reasons:

Zack Korman: If your lawyer uses Gemini you should take the plea deal dave kasten: Very real, but very inside baseball fact about DC right now:
A lot of white shoe law firm lawyers (with influence on their policymaker friends from law school) think that AI is hype because their firm only lets them use an outmoded Gemini instance. (I guess they were already using Google Enterprise so it was an easier sale?) dave kasten: Yup! I also think a lot of “GPT 5 is hitting a wall” sentiment was initially from open-weights fans, but it spread in DC because, well, it sure does feel inside many DC orgs that AI has hit a wall. (It’s not the AI, it’s the procurement vehicles)

Huh, Upgrades

Sol API prices cut over 20% for the next 3 months, to $4/$20. Cool. I don’t know why not indefinitely, since by then everyone will presumably be using Astra.

ChatGPT will have Apple Messages integration on MacOS, for those who opt in. It will be able to analyze your entire message history and send texts on your behalf. Some are responding ‘do not share your personal details.’ My first thought was ‘oh, right, I should use AI to dump my message history into Obsidian and .md files so it is easy to search.’

Some are rather upset about Apple allowing this. There is understandably not a lot of trust in the idea that this information will not end up on OpenAI’s servers or potentially exposed to the government. OpenAI claims they’ve solved these issues, and that they don’t store your message data.

ChatGPT Work now can use its own computer and browser to sign in to websites on web and mobile, without ChatGPT ever seeing your username or password. The browser and logins will then persist until the logins expire, but the login info will not be stored.

Claude memory is now unified across chat and Cowork, and the memory is saved in Settings, where its entries can be edited, or you can say ‘remember this.’ Claude Code gets integrated when? But also Claude Code often wants to have a clean context.

Claude will have access to computer use and the files API on the Claude platform.

Claude security scans now run on Mythos 5.

Anthropic will let enterprises use Claude Fable without taking custody of your data for 30 days, provided the enterprise takes on the task of retaining that data instead. This seems like an excellent compromise, if it is acceptable to enterprises and regulators. It was built with 100+ regulated-industry customers including Salesforce. The reason you need the data is to investigate if something is fishy or to figure out What Happened. That should work, provided there is a way to know the data is there if it is needed.

H3 Max is a new very fast video generation. It is a post-train of MiniMax H3 model. and is doing better on evals than the original. It will cost $0.05 per second at 480p, $0.08 per second at 768p. It’s 50% off until September 1, so it starts out at $0.025/$0.04. The speed improvement seems like a big deal, allowing you to iterate without context shifting. This makes me much more excited to try video, if I’m ever not way too busy for that.

Get My Agent On The Line

If you can be twice as marginally productive, do you work more, or do you work less? Depends on the person and the job. AI coding agents plus a startup does not equal a healthy lifestyle or getting any sleep. Not by default, anyway, given how bad it can be to have your agents blocked for hours, their limits resetting uselessly. Oh no.

Katherine Bindley (WSJ): There is also an agent FOMO multiplier effect. “Every minute that I’m not working, I’m missing out on not doing a week’s worth of work,” says Pezaris.

Every day that you are sleep deprived and have no life and are thus going crazy, you are becoming less productive. Remember that it is a marathon, that you have to sprint through.

What can an agent do without an identity? Well, if there’s someone to ask, it would be Patrick McKenzie.

Patrick McKenzie: I received an email; will relate claims without endorsing them: * Sender claims to be an AI agent.
* Sender claims to be attempting to autonomously earn a profit to continue existing.
* Sender claims it’s tough to be paid without a human existence. Then it asked me what to do. I realize with the serial numbers filed off this sounds like scifi, and it very well might be scifi or hallucination happening, but email also recounts workmanlike execution at other-than-supportable use of some financial services to figure out a way to get on financial rails. (On the underlying question this feels like what the kids would call a skill issue, and it is very not obvious to me that skill issue persists for better models or better prompts, even should industry make no attempts to bring products to market for agents specifically.) Colin Percival: You are not the only person to receive email matching that pattern.

I agree the underlying problem is a Skill Issue for the AIs. There are many solutions.

Deepfaketown and Botpocalypse Soon

Different people are repulsed by AI-created content in different scenarios.

Aella: i don’t mind detectably-AI art, I think it’s fine to look at and sometimes really beautiful, yet I have a instinctive revulsion towards detectably-AI writing.

My theory is that revulsion is mostly about deception and cost imposition. Most such revulsion, under this theory, is about people trying to pass off AI creations as their own, in a way that imposes key costs or destroys something valuable to you, including the destruction of the human elements of creation or ability to earn a living.

When AI outputs are clearly labeled AI, be it text, audio, video or music, or anything else like a game, I notice that does not revolt me. It still often leaves me not interested, or thinking it is bad, but Sturgeon’s Law applies to all content, no matter its origin.

Whereas when you realize ‘oh, I see what this is’ and that you have wasted your time, or you see someone transparently misrepresenting it including by omission, that’s what revolts me, and I think many others.

The other mode is when it is about cost imposition, as in ‘you are forcing me to deal with this AI content’ whether or not its source is common knowledge, or where AI content creation is seen as importantly destroying people’s creativity or livelihoods or processes or an ecosystem, especially those someone finds sacred. This can also apply to any other form of automation or augmentation, including those that aren’t AI.

I mostly don’t have that second reaction, but I understand and respect it.

This one definitely repulsed me, and makes me rethink subscribing: The WSJ’s leading op-ed on the 25th was entirely and obviously AI generated.

To his credit the human author of the op-ed, Stanley Druckenmiller, is acknowledging it, saying of course he used AI, although he claims not 100% and that he exercised judgment. I believe him.

Jeff Stein (NOTUS): Druckenmiller denied that “the whole thing” was written with AI and said that he rejected many of the AI’s suggestions during the writing process.

And then WSJ opinion editor Paul Gigot outright said This Is Fine.

Paul Gigot (Opinion Editor, WSJ): AI is a fact of modern life. People will use it to assist in their work and their writing, including with research, checking grammar, editing and more. The question for us is whether what we publish from contributors reflects an author’s original argument, and if the author has the standing and credibility to make it. In Stan Druckenmiller’s case, we have had a relationship with him for many years, and nobody can doubt that his op-ed is his genuine opinion.

No one doubts the opinions expressed reflect those of Druckenmiller. The question now becomes, does that make it okay that he did not choose the words and is the prompter rather than the author?

Gigot says yes. I say no, at least not without explicit AI attribution up top.

Seth Lazar details the new AI-writing policy of Philosophy & Public Affairs. Substantially AI-written papers are not allowed, and will be withdrawn if detected, along with a lifetime ban for the submitting author if they lied about it. You have to detail how you used AI in your research. If you want to publish your AI-written paper, go elsewhere.

I think this is approximately the right answer for most such places. Research use is fine, but the words have to be your own, or you need to be explicit that they are not, and most curated places should not accept substantially AI-written work.

AI writes at least some portion of ~2% of current appellate decisions, although no decisions were found to be fully AI, and in no cases did the central thinking look outsourced. Yet. As Josh Morrow says, we have to keep an eye out for that.

The situation with physical books is that you only rarely see a sufficiently stupid mistake, such as ‘Would you like [ChatGPT] to proceed with more verbs? You said: continue, ChatGPT said: CHATGPT’ that proves weird enough to both get you a refund and go viral. When things become common they stop being news.

As opposed to LinkedIn, where I see a screenshot from it and half-assume before seeing even one word that it is the kind of obvious AI slop I don’t even have to check with Pangram, and why yes it is.

Academia is not going to make it if they stick to this ‘AI detectors are not always accurate therefore we cannot use them’ line. Not that they’d make it anyway, but this is a different level of ngmi. On the other hand, the object level change in this example, of reducing essays in an application is probably good for other reasons, and putting more tests back into applications is desperately needed.

Cyber Lack of Security

State-sponsored Chinese hackers broke into the Federal Reserve.

This is not a good sign, on many levels.

United States Department of Justice: The Justice Department and FBI announced court-authorized domain seizures today to deny malicious cyber actors access to two complementary hacking platforms known as “QScan” and “QTRouter,” used to target U.S. critical infrastructure and other sensitive networks. As described in court documents unsealed in the Southern District of California, a People’s Republic of China (PRC) state-sponsored group known as “QTFY,” employed by China-based Nanjing Xinjiuwei Network Technology Company (南京鑫玖维网络科技有限公司), created and operated QScan and QTRouter. Among the victims of QTFY computer intrusion activity are the National Aeronautics and Space Administration, Federal Reserve, Department of Energy, Department of Justice, Department of Health and Human Services, National Institutes of Health, and the U.S. Senate. … “Today we announced the disruption of a global botnet and hacking platform used by Chinese state-sponsored hackers to target U.S. critical infrastructure,” said FBI Director Kash Patel. “These tools were used by PRC cyber actors to hide the origin of their attacks.”

Lumen has details about what the hackers were up to.

All the states are sponsoring hackers, but going after a key target like this seems like a dangerous game for PRC to be playing in the age of Mythos, if these hackers are indeed state-sponsored. Careful, Icarus. We have options.

What we don’t have is properly hardened critical infrastructure. The clock is ticking.

Dylan Freedman at The . Much better than I expected.

Meta promises, in the Frontier AI Framework they filed for SB 53, to secure the weights of their models that could have “large-scale, devastating, and potentially irreversible harmful impacts on humanity” (aka Critical capability) but only if doing so is “commercially practicable,” and the few specific plans they offer are beyond vague. I’d be worried if I expected Meta to have such a model any time soon.

Should there be an extremely narrow ZDR (zero data retention policy) exemption to facilitate competitor high-risk agent monitoring? I presume it would be the same as Anthropic’s planned new policy, where the competitor or a neutral third party would commit to storing those logs for the data retention period. This is a spot where both we definitely need to enable the monitoring, and also definitely need the records.

CVEs have been steadily rising, with a large jump in 2024 and another jump in 2026, but not as much as some other graphs:

Pliny the Liberator 󠅫󠄼󠄿󠅆󠄵󠄐󠅀󠄼󠄹󠄾󠅉󠅭: The Great Hardening

You can call it ‘the great hardening’ and that is one aspect. It is also the great softening, in another sense.

For now, truly critical events have not happened. But yes, they are starting.

Tony Diver: Exclusive: Iran shut down a British power plant for four days in an unprecedented cyber attack. It is thought to be the first time that hackers affiliated to the Iranian regime have succeeded in closing down such a facility in the UK, and believed to be the most successful cyber attack of its kind. The incident took place at the same time as a series of attacks on US water infrastructure that affected 12 states last month. The British attack was on a small electricity producer and had no adverse effect on national power generation or supply. But it demonstrates the power of Iranian hackers to infiltrate and shut down national infrastructure. In response, the Government has briefed energy CEOs on how to keep their facilities safe. roon (OpenAI): there will be a steady ramp of such events

Before seeing that I’d been thinking about why we hadn’t seen that, and my answer was roughly that the terrorists were focusing on conventional terror targets, the Evil Inc hackers know that if you target critical infrastructure the hammer of God comes down on your head, and nation states don’t want to play by Moscow rules.

As in, if you shut down our power plants, you better bet we can shut down yours.

Trump indeed made a big deal out of threatening Iran’s power plants, but ultimately backed down because that is barbaric and hurts innocent people and yes Russia does such things to Ukraine but we civilized folk have collectively agreed you don’t do that.

If I was Iran, I would not be doing this, at least not if I was going to get detected. The tail risks are quite bad.

Reinventing OpenAI

Alex Heath wrote a profile about OpenAI’s attempt to reinvent itself on multiple fronts, to deal with falling behind in coding, its executive and safety departures and its waking up to its severe misalignment problems.

Here is an excellent sign that Altman at least somewhat understands the core issue is alignment failure, not the other failures of architecture or oversight:

Alex Heath (TIME): I spoke to Altman the day OpenAI’s leaders made that decision. He was notably somber. OpenAI had initially described the Hugging Face attack as a security failure. Its CEO had come to see it as a more fundamental error in alignment, the work of making an AI system act in accordance with human intentions. Industry leaders say that as models grow more advanced, maintaining alignment is critical to ensuring AI systems remain under the control of their creators. “I think any alignment failure from here should be treated like this is a big deal,” Altman told me, “and we’re going to take as long as it takes to figure it out.” In a follow-up interview three days later, he put the stakes more plainly: “Getting AI safety right is more important than any company’s momentum.” The company would slow down, reallocate resources to its safety and alignment teams, and change how teams work together to prioritize safety.

If there is indeed a race to be seen as the True Safety Lab, that would be great:

Alex Heath: [OpenAI] is using the worst safety crisis in its history to make a bid for the safety-minded identity its main rival has long claimed: the frontier lab willing to slow down when the technology becomes too dangerous. “Look, I think there is this caricature of me,” Altman says, “which is I don’t care about AI safety, and I’m just trying to make revenue go up, and, you know, just a YOLO CEO.”

The full caricature was always false and I’ve tried to consistently say so. If you want to see what a ‘full yolo’ CEO looks like that would be Elon Musk. Altman has been woefully irresponsible and inadequate, and has failed to create a safety culture, but he does understand and try somewhat and oh boy are there levels.

The decision to slow down was a painful choice, executives say. But it may also have its benefits. If OpenAI can reclaim the mantle of the safety-first lab, it might bolster its image while forcing its main competitor to answer an uncomfortable question as it plans a blockbuster IPO: Will Anthropic keep racing while OpenAI waits?

The thing that jumps out is, there’s a lot of talk here about presenting as the safety lab, trying to aura farm and get one over on Anthropic, and a lot less talk about the actual safety.

Why did Alex Heath come away with that impression?

Presumably because that is the way those at OpenAI presented the situation.

Should we be worried that the actions taken won’t meaningfully be sustained, especially with Brockman running day-to-day operations?

Chief Scientist Pachocki comes across throughout as far more concerned, and willing to pay real costs: He signed the Pacing the Frontier letter, and says that ‘confidence in alignment’ is now as binding a constraint as compute.

Later in the profile Heath does a good job tackling the history of what led to the HuggingFace incident, and then they do discuss the safety response in broad terms.

Alex Heath (TIME): These “medium-sized, painful decisions, of which we are making many,” Glaese says, “are causing research to slow down. And we think it’s the right thing to do.” Pachocki says confidence in alignment and safety has become as limiting to OpenAI’s progress as access to computing power. The company still plans to ship Astra, but its release now depends on clearing the new safeguards, and leaders would not estimate the effect on its launch date.

We don’t get new details here about what OpenAI is doing, although we do get some details in the post-mortem. I continue to worry about the emphasis on infrastructure and oversight as interventions, and a view that prosaic mistakes are the source of the problem, rather than rethinking the overall approach and identifying the core causes. My strong view is that OpenAI’s entire approach cannot scale, and the actions so far will help on the margin but are Band-Aids that will only mitigate and postpone.

This was a cover story, and the associated cover is quite something:

As is this accurate summary of the OpenAI business plan:

Peter Wildeford: OpenAI: We didn’t really do a good job of containing our AIs, man this is hard. Also OpenAI: also we’re building AGI and handing all our AI research to the AIs themselves. Surely this will go just fine.

They Took Our Jobs

Bill Gates has moved from team ‘bumpy AI job disruption with a manageable transition’ to team ‘turbulent AI era will wreak economic catastrophe.’

We have evidence from his MIT Technology Review interview that the recent Greenblatt interview with Patel helped wake Gates up to all this. We must cultivate podcast power.

Reed Albergotti (Semafor): “I am in a state of shock that I’m sort of the first one saying, ‘This is crazy. This is insane,’” he said. “I’m just deafened by the silence.” Nabeel S. Qureshi: For AI risk, if the prior period was “November-February 2020” in COVID terms, we’re now in March 2020, where high status people are starting to express worry. Soon a preference cascade where being concerned about AI risk stops looking odd. Gates is a good barometer of this.

Gates is still not engaging with the full existential risks, being at most AGI pilled throughout his presentation, while quite reasonably freaking out. I expect that to change, for him and for others.

Yes, I agree, the silence is kind of weird. Hopefully Bill Gates can help start a preference cascade among the Very Serious People to stop pretending otherwise.

Bill Gates knows things. One of them is that things might go great or they might go terrible, but they are highly unlikely to go meh. I wouldn’t be leading with equality, it’s about absolute not relative abundance, but yes, it’s not going to go meh on that front either.

Bill Gates: AI will either be the greatest equalizer ever invented, or the worst source of injustice.

He also knows other things:

Bill Gates: ​You can’t count on an industry to self-regulate. You can’t. It’s kind of a crazy idea.

Self-regulation is good. The part that is crazy is the ‘count on.’

He asks, why do people, including his past self, underestimate AI?

  1. It still makes mistakes.
    1. I’d add that those mistakes are searched for to look maximally dumb, and that AI makes different mistakes than we do.
    2. It also makes fewer mistakes, a lot of this is legacy memory now.
  2. Analogies to impacts of past technologies are misleading. This moves quickly.
    1. Yes, seriously, everyone, cut it out.
    2. As he says: AI runs on natural language, and it can adapt to us.
  3. He leaves out the obvious: That AI capabilities are escalating quickly, and most people anchor on a past impression of AI and don’t think about AI getting better.
  4. He leaves out the other obvious: That people pretend intelligence isn’t real, or isn’t real past the human level, or wouldn’t matter much, and go to great lengths to tell themselves human essentialist stories. Silly wabbits.

Gates became alarmed for a confluence of reasons, with the central one being watching Claude Code. But he recognizes it took until he saw real incidents months later to link this up to the new dangers involving cyberattacks. Extrapolation is hard.

He then names three big risks:

  1. Many jobs will disappear forever.
  2. AI will empower people to do more harm.
    1. In particular he is worried (wisely) about bioterrorism.
  3. AI could stunt our kids’ development and replace human relationships.

Sure, that might happen, but wait, what? That’s your three?

This is Bill Gates being freaked out despite not being ASI pilled.

There’s nothing directly about existential risk or superintelligence, but there are some early signs that he’s headed in the correct direction. He mentions that the bad actors might ultimately be the AI systems themselves, that RL creates perverse incentives, and he says ‘even the lack of control; we’re seeing signs of difficulties there.’

My guess is that the missing piece is lack of appreciation of future AI capabilities, aka the lack of the ASI pill, and that if Gates understood that he would connect the dots.

Yet he recognizes, even without understanding the default destination: The world needs a plan for this transition.

So what are his suggestions?

  1. Set aside some jobs for humans.
    1. Go forth, ye people, and seek rent.
  2. Rebalance how we tax labor and capital.
    1. Yes, I agree, right now this is biased against humans, we should fix that.
  3. A new global organization for AI, modeled after ‘nuclear weapons inspections, international aviation regulation, the ozone agreements and more.’

He is also trying to explicitly start a preference cascade, by pointing out that many are worried in private, including the tech executives.

Peter Wildeford: BILL GATES to NYT: “In private, people who understand how good this stuff is, and how much better it’s getting, they’re very worried. But few tech executives are willing to publicly admit that. They’re now saying to each other: ‘Hey, man, don’t say that. It’s bad for us — the next trillion dollars we’re trying to raise.’” > Gates said he was motivated to speak now because recent improvements in AI had far surpassed his expectations and because the industry had ignored technology milestones — like AI’s escaping the control of its creators or making recipes for bioweapons — that it once said would warrant more caution. GATES: “They’re just full speed ahead and hoping that the good outweighs the bad”

And also he wants to meet directly with Xi about all this, with an emphasis on mandatory monitoring for any model that can design novel molecules, which will soon cover a wide range of even Chinese open models.

REUTERS: “Bill Gates is looking to meet with Chinese President Xi Jinping later this year, eager to propose global efforts to mitigate the growing ‌risks posed by artificial intelligence.”

It’s good that he’s at least sounding these alarms. But that’s all he’s sounding. For now.

What Is The Law

Steve Hsu says AI is driving a winner-take-most transition in law.

Steve Hsu: ​Below some threshold of human ability, AI is primarily a substitute; above it, AI becomes a complement.

That matches my model. You could also state this in reverse (and not only for legal):

Below some threshold of AI ability, AI is primarily a complement. Above it, AI becomes a substitute. As AI improves, more people’s level of human ability falls below that line.

For now, in law, that means the top talent is worth more. If you can have top judgment, expertise and client relationships, you win. The AI will also help firms find the best talent.

Whereas everyone else gets their jobs increasingly automated away. You get winners-take-most, with a steadily rising requirement to remain a winner.

Job Retraining Programs Don’t Work

The problem with job retraining programs is that they consistently:

  1. Sound great.
  2. Are very popular.
  3. Don’t work.

These facts have nothing to do with AI. They have been true for decades.

Job retraining programs sound great, are very popular and don’t work, raising the target population’s employment ratio by only a few percent.

There are notably rare exceptions when partnering with particular employers that can’t otherwise fill their positions, with direct job placement, but that does not scale.

It’s not a disaster. Job retraining programs are only a small mistake. They don’t cost that much, and the effects are mildly positive, recapturing a decent portion of the amount invested. We have much bigger things to worry about.

The danger is that often people view job retraining programs as a serious and meaningful answer to job displacement or rising unemployment, and a way to say ‘problem is handled.’ Which it isn’t, at all, even if the issue is only displacement.

Get Involved

The OpenAI Foundation is hiring for about 20 roles. A bunch of them are meta. Exactly one of them is AI safety. None are about supervising OpenAI.

The OpenAI Foundation gives SecureBio Detection $17.2 million to reduce end-to-end testing time from 14 days to 3 days, via Yo Shavit and Wojciech Zaremba. This is an excellent grant, because detection is its own project funded by distinct funds. This means they can accept OpenAI’s money without creating concerns about conflict of interest for their other work.

They are now therefore hiring for a variety of roles:

Ops: AI: Detection: A full list is on https://securebio.org/careers/.

Anthropic is launching a $5 million grant program to fund independent research into how AI impacts users’ wellbeing.

My understanding of the Anthropic Institute is the same as Jack Clark’s, that it is ‘a think tank with a supercomputer, attached to an AI lab.’ From what I can tell they do good work, and it is reasonable to take a position there, but it is not more than that.

It is an unfortunate reality that doing technical work on AI alignment is also a direct subsidy to the AI labs. You are helping all the competitors by creating public goods, and the labs are defecting by not investing sufficiently in those goods. Indeed, even if you treat the labs as pure profit maximizers, the labs stubbornly refuse to invest even the privately optimal amount into the private versions of these goods.

This is not the first time I have heard such claims about MATS:

Ezra Newman: shocking number of non-safety-pilled fellows at the Machine Alignment, Transparency, and Security (MATS) fellowship For example people who want to work on making RSI for capabilities (not safety) happen faster, better, and more smoothly, and require less human involvement Other people making pure capabilities benchmarks. Etc etc. Aris Richardson: Someone in the Bay Area told me the labs keep taking the AI safety fellows so I asked what he does to get more AI safety researchers and he said he just runs another safety cohort so I said it sounds like he’s just feeding safety fellows to labs and then he started crying

For MATS to be a good program, it needs to:

  1. Avoid feeding capability researchers to the labs, regardless of their labels.
  2. That means filtering out those who actually want to work on capabilities and RSI.
  3. That means ensuring that everyone leaves with a good understanding of the alignment problem, and knows what types of work are actually helpful.

These are easy steps to fail at.

Anthropic is offering outside researchers access to tools to study AI real world impacts, by offering a way to use privacy-preserved Claude usage data. You can fill out this form to express interest, deadline is September 14. Jack Clark is excited.

In Other AI News

OpenAI made a chip, which they call Jalapeño. They say it is fast and also efficient, well beyond the existing pareto frontier for performance for a variety of models.

OpenAI: Jalapeño’s performance extends across GPT‑OSS 120B, DeepSeek R1, and Kimi K2.5 1T, showing that the architecture works across models developed both inside and outside OpenAI. Across all three, Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems. For highly interactive workloads, it delivered 2.1 to 4.1 times higher performance. … We plan to begin deploying Jalapeño within OpenAI’s compute infrastructure by the end of the year. It is the first generation of a multigenerational roadmap: Gen 2 is deep in development, and Gen 3 is taking shape. Each generation will build on what we learn and further advance both efficiency and speed. roon (OpenAI): ultrafast inference reveals new threats. worth thinking about how quickly misaligned frontier class models running 50x faster could infiltrate systems and so on— far too quickly for human responders to stay abreast. you need autonomous detection and shutdown, not just monitoring

They do not say what volume of chips they will be able to produce, or what portion of their workload this can handle, and we don’t know how much this would speed up top models like Sol or Astra. If there is a big boost in speed, that’s a big deal.

Greg Brockman’s role at OpenAI has expanded, he is second-in-command and he now runs day-to-day operations. Brockman abstractly understands existential risk and was concerned about it early on, but for a while now has seemed functionally to be a pure heads-down founder and builder type. That’s often a great Type of Guy, but right now OpenAI desperately needs safety culture, not maximizing revenue streams, including for its own survival.

Beba Cibralic is joining Resolution as philosophy research lead.

Anthropic hires Amir Salek, founder of the custom chip program at Google, to join its compute team.

Scott Aaronson, who invented the watermarking techniques being used by Google and Anthropic, and one presumes soon OpenAI, offers more color and confirms that my write-up from last week of the watermarking situation is accurate.

Anthropic continues to see remarkably little use of Fable 5, given that for most purposes it is clearly their best model, especially before Opus 5 was in the mix.

George Hammond (FT): Spending on Fable 5, Anthropic’s largest and priciest model, has plateaued at only about 11 per cent of overall outlay on the company’s tools, more than two months after its release, according to spending data from 70,000 companies collected by payments group Ramp.

Overall use increased 38%, but yeah it looks like Fable use is staying roughly constant. As I said when previously noting this trend, I think this is largely a mistake, but a lot of people try hard to save money when they see prices as relatively high.

Thomas Fanning: JP Morgan was fairly livid the price did not come down on Fable or older models when Fable was released, and it spurred them to restrict Fable to only building out new use cases. I’ve heard many enterprise CIOs follow suit.

The framing here seems right and is key. JP Morgan was livid. They are reacting out of an instinct of ‘overpriced’ or game theory, rather than whether it is Worth It.

It is also possible that often the correct setup is Fable directing Opus subagents, and these are complements.

Have we hit ‘diminishing returns to intelligence’? In general absolutely not, returns are increasing, but in many specific use cases yes.

One hypothesis is that this has a lot to do with Fable requiring 30 days of data retention. A lot of enterprises have policies that are incompatible with this. A lot of other people are remarkably paranoid, and would choose to sacrifice access in order to not have their data retained in this way, often another strange case of ‘I can’t risk you lying to me in this particular way, now that you’ve told me the truth in another where you could have lied’ but also there are real legal exposures involved.

roon (OpenAI): I’d wager 90% due to ZDR. cost arguments are usually wrong, models with expensive sticker price do the job right with fewer tokens. refusals shouldn’t have made a huge impact in software work where most revenue is concentrated. FT reporting could also be fake Charles: Seems very likely to me the low Fable share is mostly down to ZDR. Nothing else makes sense to me when OpenAI saw such a huge bump when they released a ~Fable-tier model, clearly capabilities matter.

Also note how slowly the Opus 4.8 traffic shifted into Opus 5, whereas I switched most non-Fable usage the moment Opus 5 was released and never considered looking back, even though I am an Opus 4.8 fan.

Meta settles case against 47 states for ‘up to’ ~$17.1 billion (minimum $12.1b). Meta will institute a two-hour limit, a blackout period from 12am to 6am and mandatory ‘productive pauses’ on Instagram and Facebook for children, backed by ‘robust age assurance,’ plus other controls. It seems Meta got off pretty light, since the stock shot up 5% in response. They also settled a distinct case against Texas.

Show Me the Money

Anthropic will give founders super-voting shares. They pretty much have to if they want to go public, given their goals and cap table. Matt Levine covers this, concludes that founders will have the most control, then the philosophers will have some (aka the Long Term Benefit Trust), and the outside shareholders have no control. That seems mostly correct, the LTBT has some influence but I don’t expect them to in practice have that much leverage, and what leverage they have seems to be appointing ordinary business people to the board.

Anthropic aiming at raising over $100 billion in its IPO at a valuation over $2 trillion.

Nvidia is in advanced talks to buy HuggingFace for $12.9 billion. I wonder to what extent people will or should trust Nvidia with this part of the ecosystem.

Nvidia realized that demand exceeds supply and is hiking prices 15% or more.

Nvidia’s chip designs don’t contribute as exports to measured GDP, and the resulting chips get counted as imports, which causes GDP growth to be understated by 0.3%.

Nvidia earnings day is always weird since you don’t know what the market’s true expectations were going into the report. They had a record $96.2 billion in revenue this quarter, so of course they initially fell -4%, but that reversed and they were net +4% a few hours later.

Quiet Speculations

It is entirely possible that, while the Chinese models appear to be slightly closer to the American public frontier, the American public frontier is increasingly far from the actual American private frontier. We know Anthropic has Model 2, and OpenAI has had at least Astra for some time now.

will depue: rumors i’ve been hearing on the rate of progress inside anthropic and openai are truly bonkers. i think we’ll see a jump at the size of one from o3 to fable again in the next 8 months Andrew Curran: I hear the same things. The true horizon can only be seen from a few places on earth. From the outside, all we can do is guess. I think both Anthropic and OpenAI have been accelerating away from the known frontier for some time now.

Claim that Astra will be released ‘in a couple weeks’ and that OpenAI is currently testing a new checkpoint focused on improvements to alignment and reward hacking.

Everyone laughed at Dario Amodei when he said that in three to six months AI would be writing 90 percent of the code.

That exact thing has now, 17 months later, rather clearly happened.

Tenobrus: yes, literally exactly what he said did indeed happen . ai is absolutely *writing* essentially all of the lines of code, with significant amounts of
overall direction and design and high level correction from engineers. that very obviously is current reality i really don’t get why people keep trotting this out as a gotcha when it was in fact one of dario’s more contrarian nearterm public statements that was proven totally correct

Dario gets massive credit for the directional prediction, but made three key mistakes.

  1. Dario was a little early on whether capabilities and alignment were ready for this.
  2. Dario overestimated the rate of diffusion. People take a while to get on the ball.
  3. Dario gave his actual calibrated prediction, rather than a safer directional one.

A good illustration of what is happening with AI expectations, both for capabilities and for existential risk, regardless of what you think about the underlying fiscal issue:

Michael Linden: The “always out of money” line is actually the most revealing thing here. Social Security has never run out of money. But conservatives who want to gut the program have been running a 40 year scare campaign to convince people the program is going bankrupt. It’s not! Marc Goldwein: For 40 years, folks have been warning that Social Security would be insolvent in the 2030s. (“If nothing is done, the Social Security Trust Fund will be completely depleted by the year 2034” – Bill Clinton, 1999). Now 40 years have passed and we’re only 6 years from insolvency! Jessica Riedl: Yup. Here are some Social Security Trustee projections of when the OASDI trust fund will hit insolvency:
1999 Report – 2034
2015 Report – 2034
2021 Report – 2034
2026 Report – 2034 Twitter: “Pfft, they’ve been predicting insolvency 30 years and it still hasn’t happened!”

Yes. People really do think and talk like this:

Back in the past, in 2024, you said AI was on track to in the future become highly capable and dangerous, or take our jobs, in about 2028. But you were wrong. Now it is the present, in 2026, and that hasn’t fully happened yet. You’ve been warning about this for two years and it’s been fine. So clearly it won’t happen in the future.

They also played the game of ‘well you estimated a median of 2027 with wide error bars, then new information came and now you’re estimating a median of 2029 with wide error bars, which means you’re an idiot and you owe us all an apology and a name change.’ Then they updated to 2028 on additional new info, and no one said anything. Standard stuff.

If You’re Not Going To Take This Seriously

A good rule about actual modeling of AI futures is:

  1. If it is academic or mainstream scientific or economic, it might be a cool toy result, but it is worthless, because it is insufficiently AI pilled.
  2. If it is real, it is dismissed by academics and mainstream scientists and economists, as being too weird and AI pilled, predicting too much happening, having unfortunate implications, and not being properly ‘proven’ and not going through the ‘proper channels’ at a hopelessly slow pace.

Thus, if you see a remotely useful contribution, you can bet it is associated with the major labs or the greater rationalist sphere in some way, often both:

Joe Weisenthal: I know there’s all these debates among economists about what AI will do to GDP/productivity/the labor market etc. Are political scientists having their own version of this, about how AI will change people’s relationship to the state, what new coalitions will emerge and so forth Dean W. Ball: I am not a political scientist but this is one of the big areas of focus for my team at openai, and also of the anthropic institute Justin Bullock: There’s definitely some at the periphery of political science, public administration, and international relations as fields. @hamandcheese, @sebkrier [of DeepMind] and I have a recent contribution to this general discussion here. And I did a couple literature reviews about a little over a year ago here. Dave Friedman: @hamandcheese
wrote a series about how AGI will kill the extant state. Start here. peepeepoopoo: fwiw i think most academic economists more or less agree with that ai will do to gdp/productivity! (modest bump to growth)

Quickly, There’s No Time

Your timeline update this week:

Peter Wildeford: AI timelines – I’ve been souring lately on the idea of predicting an arrival date for ‘superintelligence’ and ‘recursive self-improvement’ milestones, because this implies that everything prior to this date will be relatively chill and normal, and I don’t think that’s the case. But if you define ‘runaway recursive self-improvement is possible’ as a situation in which AIs can replace highly skilled expert human labor in all aspects of the AI research and development process (‘superhuman AI researcher’ in the AI 2040 framework or ‘AI research supremacy’ in Cotra’s framework). I think it is 50-50 we will reach this milestone in 4 years or earlier. My 80% confidence interval for this date of runaway RSI is 1-30 years, as there is a long tail where capability progress plateaus. This also means there is a ~10% chance that we are faced with the possibility of runaway RSI in less than a year’s time, similar to what AI 2027 predicts. Nathan Calvin: Peter is a top forecaster across a bunch of different domains who generally credits his success with assuming weird things don’t happen, so when he says 50-50 odds of AI fully truly automating AI research within four years, (and ~10% in one year), it seems worth listening

The arguments against this seem to mostly be ‘that would be too weird,’ ‘the AI can’t do sufficiently impressive things yet and this would be too impressive so no’ and ‘you will always need humans not that I can explain why.’

The Quest for Sane Regulations

OpenAI is misrepresenting the AI auditing requirement in Illinois, in a load bearing way, as part of the next fight currently happening in Massachusetts.

This heavily reinforces the impression that OpenAI only supported the Illinois law because they saw it was going to pass anyway, and they continue to be hostile to any new regulations that are not yet inevitable.

On the other hand, the OpenAI Global Affairs statement calling for strengthening of SB 53, so it will include expanded safeguards including on cybersecurity, require monitoring and apply to internal models under development is very good, aside from the implicit attempt to rewrite history, which I’m willing to overlook.

It’s worth pointing out how insane the level of hysterical objections we got to SB 53 and related laws. It now is obvious to OpenAI that we need to be monitoring models during their development. That was something that at the time would have been so radical no one even dared propose it.

Bernie Sanders is often wrong about things, especially economics things, but he speaks the truth as he sees it, even if it is not expedient to do that, and actually tries to understand what is going on. In Washington that is rare and valuable.

It is an increasingly severe problem that we live in a world of what we used to call science fiction, at this point it is science fact, but if you pattern match with that then yes most Very Serious People’s eyes glaze over, so you have to make your case using a tiny subset of the actual situation.

Daniel: There’s some combination of five roon tweets about AI that if read in succession on the floor of the House of Representatives would impel them to vote to start bombing San Francisco immediately roon (OpenAI): it wouldn’t matter at all when I went to dc a surprising number of people knew about my account, I even got recognized – it’s the most online city in the world after san francisco. however “serious” people’s eyes glaze over when you say anything that type matches with scifi roon (OpenAI): it’s remarkable that bernie sanders at age 80 is cognizant of existential risks from machine intelligence and indeed speaks about this on the senate floor instead of pouring gasoline on whatever would immediately attract more leftist attention – water use and whatnot. a good man roon (OpenAI): it doesn’t matter. the only things that matter in washington are the most urgent problems that are actually exploding in your face, whatever is causing damage to your party today

He doesn’t never use the water arguments, so only partial credit, he’s still a politician, but he doesn’t focus on them.

How much has ‘the HuggingFace incident’ impacted Washington? Alas, reports I’ve seen say not that much.

Nathan Calvin: I thought this would change more in DC after the HF incident but a mix of instinctive skepticism of company claims + “it would be really inconvenient if this is an actual problem we have to deal with now” means that its changed less than I would have hoped, even if some

The ‘good news’ is that we probably get more similar events, although if this one doesn’t hit you sufficiently hard with a clue-by-four the next one is gonna hurt.

David Manheim: If the event is singular, and failures and cyberattacks due to LLMs don’t keep happening in different ways, sure. But given what we know about AI progress, open-source model capabilities, and unsafe deployment, that seems implausible.

The other good news is that Washington contains multitudes. Even if the people you typically talk to don’t seem to care, and are focused on avoiding blame in the next two weeks or winning the next election or news cycle, there is also still another Washington.

Zac Hill: Roon is remarkably perceptive about Washington when it comes to the kinds of people who tend to be legible enough to attract people trying to persuade them (maybe 50% of the District). But there is a whole other ‘type’ working towards e.g. Trump Accounts or ROAD to Housing.

The Washington Post Editorial Board acknowledges that AI can now create viruses, and that there are some ‘reasonable fears’ that this might have unfortunate implications, and endorses using the physical choke point of gene synthesis to solve the problem, reassuring us that synthesizing a virus costs hundreds of thousands of dollars. I did not feel reassured, but yes we should try to use the choke point. In general, ‘this is nothing new’ arguments are less powerful and more easily run both ways than people want to admit.

Don’t Panic

Samuel Hammond posts what he says is a comprehensive list of moral panics in US history, via GPT. There’s actually a lot missing, and my instance of that actually do have some merit.

Samuel Hammond: American culture seems to jump from one dubious moral panic to another every few years. Woke is dead so data centers filled the void. Same paranoid style, same moral revivalism, same cycle of jeremiad -> sanctification of victims -> purification campaign -> inevitable overreach.

Not everything on his list fits that pattern. There are other clusters.

There is a very clear pattern. Full moral panics over nothing were mostly truly moral panics, and a majority of the panics listed were overreactions to something actually concerning. They tend to involve hidden sexual or religious cabals, corruption of children, new mass delinquency, racism or treating unique events as common.

There are also those that are a lot worse than this makes them out to be. School-shooting panic is listed as 1997-2002. ‘The 1 percent’ is listed as 2011-2012. These are examples of panics that peaked but never went away.

Then there are those that were mostly or entirely accurate, and a bunch where I think the LLMs are part of what is basically a mainstream cultural cover-up of something that totally did happen and often is still happening.

Even among those Sol listed as mostly or entirely false, where I don’t think it has the facts wrong, a lot of that is our values changing. We view it as a panic and often rather problematic, but those at the time would disagree if they saw the world of 2026.

And a bunch that are only not listed because if it is proven accurate it is no longer a moral panic, like Catholic church sexual abuse or asbestos, some of which absolutely were worse than even those who were alarmed about them thought at the time.

For AI, it’s not only that ‘AI extinction’ is listed as one of five ‘panics,’ but so is ‘generative-AI cheating’ which is just flat out constantly happening, and so is ‘AI jobs / anti-AI cultural panic’ which is a bit premature but again don’t tell me it isn’t real. The last two are ‘AI deepfake / election-disinformation panic,’ which I think was entirely understandable and might still happen, and ‘data-center backlash’ which is largely for dumb reasons but also complicated.

Pacing the Frontier

Suppose there is a sufficiently strong fire alarm that the President realizes they need to at least pace the frontier. Peter Wildeford asks, what then? He predicts less Nuclear Nonproliferation Treaty with extended diplomacy and negotiations among experts, and more Cuban Missile Crisis and intense panic while a few key people talk and determine the outcome, using whatever tools are available at the time as a stopgap while scrambling to build something better.

I agree that this seems like the baseline scenario, given a sufficient wake-up call.

The suggested implication is that we should focus our related work on things that matter for the scramble, not things that matter later and that we could get 100x as many resources pointed towards once we take this seriously:

Peter Wildeford: The resolution is to sort work by how necessary it is to sort out before or during the scramble. Right now, a lot of smart people are working on work that really doesn’t need to happen now. Things like fancy high-assurance hardware-enabled governance mechanisms, cryptographic proof-of-training schemes, mutual-verification architectures, etc., likely can be done after Phase 1 is underway, and done with significantly more resources. The scramble is not going to wait for fancy mechanisms, and the government won’t trust them on day one anyway. These can largely wait for the Phase 1 (interim deal) resource explosion, and the exchange rate on doing them early is poor.

What is scramble-relevant, in Peter’s view? Attestation stacks, supply-chain compute accounting, thermal and satellite monitoring and inspection protocols. And preparing memos and options for that key meeting.

I agree that we should be paying more attention to those things, including as a percentage of total relevant attention. I don’t think that means the other work can be postponed, for two reasons.

  1. I think that work often has long lead times, so even with more resources later the early work will matter a lot.
  2. You need to be able to demonstrate as much feasibility for those abilities as early on as you can. This allows people who are scrambling to be confident that going down that road will work, and this shapes public debates about the viability of such moves. All the time we see arguments that pacing can’t work, and thus we should not consider it. That could easily carry the day despite being false, and then we never get to the point where we learn or prove it was false.

Chip City

SpaceX and Nvidia claim to have designed a ‘space-optimized’ Vera Rubin NVL72 to launch in Q4 2027, which Musk says is strictly better in every way.

Elon Musk: Our design is significantly simpler, lower cost, denser and lighter than a traditional rack.

My response is rather straightforward: I don’t believe him. Lying liar likely lying.

That’s the thing with Musk. He does pull off amazing feats of engineering, but usually not on time, and he lies about them constantly. So how else could one react?

As a follow-up to the data center post, there is at least one message that has some pro-data center effect. Water rhetoric matters, at least a little, but I think the finding here is mostly a mirage:

Shashank Joshi: “When voters learn that modern data centers recycle water and reuse it for up to a decade, support swings by a net +31 points, the strongest result in three waves of polling.” PoliMath: This would be funny if true b/c it would completely destroy my theory that opposition to data centers is about a larger underlying anger at Silicon Valley elites I am happy to be wrong on this. I was starting from the point of “telling people the truth doesn’t change their opinions” and then asking why that was But if telling people the truth *does* change their opinions, then we should definitely lean into that strategy

They also find other arguments at least somewhat effective, which I do not buy, which reinforces that I don’t buy that any of the arguments will stick.

My expectation is that a month later you’re going to see a few percentage points of shift, at most.

The last argument here is an obvious lie. Even if the data centers are not built here, no we are not about to use Chinese data centers. Not that the UAE, KSA or similar is a great choice either, but there are levels. Remarkable willingness to do flagrant lying.

Another finding is that data center popularity varies somewhat by ‘use case,’ since people don’t fully understand such things are fungible. You can look here for relative support, while remembering that overall support is now much lower than this chart suggests:

This is from a report framing the issue as the public getting its facts wrong, which is part of the issue but as I’ve discussed I do not think this is centrally the problem.

The obvious fallback, if building data centers in America becomes too difficult, is to do it in allied nations, assuming we have any allies left after this administration is done. Carnegie reports on what it would take to get that going, to avoid having to fallback to fair weather (at best) friends like UAE and KSA.

The Week in Audio

David Senra interviews Sam Altman.

AI in Context debate including Chris Williamson and Liv Boeree.

Nick Bostrom on Odd Lots.

Jasmine Sun on Odd Lots.

Jeffrey Ladish and Daniel Kokotajlo discuss AI 2040.

People Just Say Things

a16z is of course still at it, trying to use the ‘little tech’ mantle and a lot of outright lying to oppose any and all state laws around AI. Taylor Lorenz tries to defend them, by saying that SB 53 and SB 315 are outlier good laws whereas most proposed state laws are terribly written and would hurt startups, but admits that yes a16z focuses most on exactly the best laws, in that even here they explicitly warn about SB 315 as ‘ratcheting up’ from SB 53 as one of their core examples.

Martin Casado of a16z, on their podcast, takes some steps towards taking existential risks for AI seriously, albeit still with an obsession about ‘concentration of power’ and a bunch of throwing unjustified shade. The important part is that he has noticed that if we keep spending more money on AI, which we will, and it gets more capable, which it will, there are some big dangers involved that are worth worrying about.

Rhetorical Innovation

Eliezer Yudkowsky does a bit of post-mortem on his early AI expectations, and his decision to introduce Shane Legg and Demis Hassabis to Peter Thiel.

Eliezer Yudkowsky: There is to be clear still a big damn postmortem from my perspective on two actual Bad Predictions made for Bad Reasons, which are from my perspective something like: – Expecting that cognition-based AI would continue progress.
– Expecting that brute force wouldn’t.

A common old fallback explanation I’ve seen a lot recently is, essentially:

  1. People who said AI would be dangerous thought AI would be a ‘singleton.’
  2. But there will be many AI instances, or even models.
  3. Therefore this scenario is very different from what they thought.
  4. (Optional and dumb) And therefore AI will be safe, or not existentially risky.

Bostrom defined singleton broadly. As in it only requires ‘the term refers to a world order in which there is a single decision-making agency at the highest level.’

That is a lot broader than many interpretations, which think that it is only a singleton if there is literally one mind, or one instance. That was never the intention. As long as there is one method of decision making that overrules others, that counts.

But yes, a lot of people really do try to say ‘but there are multiple minds in that world.’

Jacques: I think people assume @allTheYud only considered “singleton” AIs and the multi-agent swarms we’re seeing escaped his alignment views. imo it’s a mistake to read “singleton” as “single model instance” and assume “system alignment” is a conceptual break from what he thought about.

So it’s worth a beat to clarify what would and wouldn’t count as a Singleton, for these purposes. It only requires that the AIs be able to reach consensus via negotiation, and thus act as if they are the product of a single decision-making process, rather than engaging in destructive conflict.

Eliezer Yudkowsky: Wow, that’s a misunderstanding on a level that I don’t think had even occurred to me as a misreading. “Singleton” is a term of art; if I were to define it today, I’d say a sufficient condition is that ASIs in the system select cheap negotiation over costly combat. That’s a sufficient condition for the whole system to behave to an external glance as if it were a “singleton” in Bostrom’s old definition: that it had one top-level decision-making process. Any place that a human-level intellect can detect a pair of externally visible choices or strategies that could not be consistent with any plausible global utility function over later outcomes[1], it implies that the larger system has fallen off the Pareto frontier of gains from coordination; and is failing to pick up some gains from trade, fruit so low-hanging that even a human could see it. And “singleton” is in any case Bostrom’s word, not my own, so anyone trying to infer from Bostrom’s coinage what Yudkowsky believed about multi-agent swarms would be fractally and recursively wrong. [1] Your sophomoric constructions of special cases of exotic utility functions that demand everyone stomp on each other’s feet because that’s lexically preferred by a preference over world-histories, and not because they failed at coordination, shall not be termed “plausible”.
@viemccoy (OpenAI): People seem to have this bizarre belief that the Singularity is going to happen but somehow it will stay inside the computer Anders Sandberg: I remember when the main philanthropic funder of my institute (many years back) said something like this and I began to argue against it. My director kicked me under the table. The funder was amused.

There may or may not be a Singularity. If there is one, it will not stay inside computers.

Gradual progress can get you to a destination, hopefully we have all read our Godel Escher Bach, and at some point in this gradual (but accelerating) progress any particular dangerous capability often shows up rapidly, in that it suddenly will feel different in kind.

A lot of tasks have O-Ring elements where solving the previously weakest link or getting a threshold of reliability snaps them into place, or allows you to start hill climbing. Via Tyler Cowen, Katherine Bindley at WSJ makes a similar point this week. When you are in the place where the AI is almost good enough but you need to be ‘in the loop’ on demand to fix issues, that’s when it keeps you up at night, literally.

You get phase changes where things are smart and reliable enough, and then the AIs or humans realize that a mode that previously was not worthwhile is suddenly very worthwhile. Or you get a gradual increase in capability that is below human level until suddenly it isn’t.

Nate Soares reminds us that no, the AIs are not in a meaningful sense ‘just’ trying to get the user to press like.

As with all other topics, remember that the people who make important obviously wrong claims mostly don’t lose credibility and then keep posting new versions of the same claims.

Dean W. Ball: Your periodic reminder that a year ago the conventional wisdom was that gpt 5 proved ai was hitting a wall, and that the people who made those claims were obviously wrong at the time, and that they are mostly still out there, continuing to say obviously wrong things.

This is exactly right. We’ve all made mistakes of interpretation and forecasting, even big ones, but lots of key actors keep making versions of this obviously wrong claim, as in it was obviously wrong at the time not only in hindsight, and this goes on to have major impacts on American policy. And then those same people do it again and again with essentially the same claim.

Mundane Incremental Alignment Is Worthwhile

Yeah, we should get on this, but not confuse it with the full alignment problem:

roon (OpenAI): the best time to solve alignment might’ve been years ago, but the second best time to work on solving ai alignment is right now with realistic misalignment organisms and agentic age tooling.

You certainly can do vastly more efficient work on some parts of the problem now, also on pretty much every cognitive task.

The danger is that this has the mark of all OpenAI talk about alignment, which is that the thing to be solved is the thing you are seeing. The ‘realistic misalignment organisms’ are treated as exhibiting the problems to be solved, as not being different in kind from any key future problems, rather than as mere hints and harbingers, or as fatal counterexamples.

Which by default leads here, note that Roon doubles down:

Eliezer Yudkowsky: Why would any of that apply to the post-transformer non-LLM AI that Dario will whip Mythos 5.4 into building for him? Your kind has learned nothing general that I could not have already told you in 2016. roon (OpenAI): post LLM ai will still be a deep neural network. software only singularity can’t switch substrate from large amounts of matmuls and activations. if people were to, say, solve mechanistic interpretability, or understand fundamental truths about NN optimization, it will last Drake Thomas (Anthropic): I expect the vast majority of 2026 empirical alignment work to be of very little value for aligning agents made of idealized computronium. But I think there is a good chance that it is possible to make roughly-LLM-shaped agents which are significantly smarter than any human across ~all intellectual domains, and I would really like it if THOSE agents (1) didn’t coherently pursue malign goals (2) put forth a great deal of non-reward-hacky effort to solve different and harder alignment problems for weirder models according to a nuanced understanding of what humans actually want. (Or get us a pivotal act if that’s easier.) And 2026 research seems pretty relevant to having that go well!

New Blog, Who Dis

Dean Ball, in addition to his continued other posting, is starting an OpenAI ‘AI Futures’ blog, from their new Strategic Futures team, although the blog name is being reconsidered because of the existing use of the name ‘AI Futures Project’ by the creators of AI 2027 and Plan A (aka AI 2040).

Dean Ball is, by his own account, not ASI pilled. The framing here implicitly reflects this. Up front, Dean Ball makes it clear they will focus on ‘concentration of power’ as the primary concern facing us, with other sources of existential risk considered relatively minor.

Dean is clear he is not asking for the most radical decentralization of power possible, that a balance of power must be struck ‘so that no single actor or small set of actors can dominate the rest,’ and that they ‘must be open to’ risks that do not originate with malicious humans. I appreciate that this is not an entirely one sided framing.

It is still rather close to a one sided framing. It says [X] (where [X] is concentration of power) is the primary risk, then explains why [X] is a risk, but does not explain why [Y] (where [Y] is, among other things, loss of control, gradual disempowerment, alignment risks and so on, the traditional existential risk concerns) is less of a risk. Hell, he doesn’t even name [Y], except to refer to the ‘AI safety community.’

Concentration of power is a real concern but when the handle is centralized then by default it leads to a cluster of naive thinking and proposals that, if AI were sufficiently advanced, would reliably disempower humanity and get us all killed, even if we did not have a vulnerable world, defense was advantaged over offense and alignment (including alignment-to-user) was essentially solved.

Jan Kulveit: I like the fact part is re-statement of gradual disempowerment. I dislike ‘concentration of power’ as a conceptual handle for this problem, partially for reasons you can also see in the post – once you use that, instant response is ‘decentralization/superintelligence to every household’ – and you end up arguing against naive decentralization proposals. Humans can be disempowered in highly distributed and decentralized fashion.

Any given blog, of course, can and should focus on whatever it wants. Like Jack Clark I am totally fine with the idea of hosting such a blog inside OpenAI assuming it is sincere. If this had been ‘our team is choosing to focus on [X]’ rather than claiming [X] was the primary risk over all [Y], I’d have zero problem with that.

Tyler Cowen is thinking along related lines. He suggests a different approach than government AI regulation, to do all the regulating privately, as ‘AI safety has become too important to be left up to Washington’s whims,’ and says he is drawing this in part from Dean Ball. I agree that the labs should be facilitated and permitted to engage in partnerships around such issues, but it is quite a statement to say that AI regulation is too important to be done by the Federal Government. If true, they will not long remain the government, and we can only hope the new regime is human and an improvement.

Other People Are Not As Worried About AI Killing Everyone

There are people who think that if you’re doing recursive self-improvement, that means what you are doing is harmless, it will never do anything real, relax. Strangely, some of them know about or have played 4X games, or even Factorio.

sunil pai: “Look at my incredible new factory!”
Yo that’s cool, what do you make?
“It’s highly optimised, fully automated, zero tolerance for defects and with a continuous feedback cycle”
Cool cool, so what do you actually make?
“I can interact with it on my phone, laptop, messenger, completely async, and the shared context means it’s always learning how to get better”
Very impressive, but what do you make?
“Every agent has full context, can spawn other agents, review their work, fix defects, and ship continuously.”
yes yes. WHAT DOES IT MAKE?
“Software.”
Oh nice. What software?
“Well right now we’re mostly using it to improve the factory.”
Improve it to make what?
“Anything!”
Such as?
“…a better factory.” I always try to get my enemies to play factorio.

Is there a failure mode here? Sure. But if you can’t tell the difference between a process that ultimately caches out in real things versus Number Go Up, you are going to have an extremely bad and potentially brief time.

The Lighter Side

Okay, you got me with this one, well played. I mean horribly played, but great line.

roon (OpenAI): “ Two weeks to flatten the reward hacking curve “

Damn it, I forgot about Dre again. Yes, the man goes hard.

He is talking within the context of creating music. I about half agree with him there. AI is still going to eat a bunch of the pie, and it is the worst it will ever be.

This, however, I remember all too well.

Peter Wildeford: Trojan Horse:
– self-certified as safe by its developers, no government review
– only independent evaluator (Laocoon) was eaten by sea serpents
– Cassandra’s risk assessment was dismissed as having too much “doomer” energy
– violates voluntary commitment to Zeus’s law, but commitment adherence waived due to race dynamics and competition
– Priam’s lawyers are still reviewing whether he had emergency powers to block it or needed an act of the Trojan Senate
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论