AI #188: Gemini Dot Argon
Is Google back?
They claim that they are back. Gemini 4 Argon is rolling out, with competitive frontier-level benchmarks, at $2/$10.
What we don’t have is access to the model, because Google Fails Marketing Forever. So it is far too early to say what we have here. When I know more, so will you.
OpenAI was forced to pull what would have been GPT-6.1 Astra due to alignment failures. They did offer us GPT-6.1 Sol, which is pitched as approaching Astra quality at the much lower price of $2/$10, the same as Gemini 4 Argon.
The rest of OpenAI’s big Dev Day announcements were Ultrafast mode and Dots, your always-on AI agent based on Astra, which comes with your Pro subscription. I’m trying it out and will report back over time if I find it useful.
The new hotness remains Claude Opus 5.5. This model rocks. It has made me considerably more productive and made my day more pleasant. It should raise your ambitions. There are some particular reasons to call upon Fable 5.1 or Astra, and sometimes a cheaper model will do, but pending Argon I consider Opus 5.5 my favorite model by a wide margin.
Politics continues. The White House hosted a lunch with the major tech leaders, and everyone signed a ‘morally binding’ White House accord on AI Safety. This gave the labs the go ahead to meet and agree on safety standards, and also called upon them to complete the Quest for Embedded Evaluators.
Jensen Huang went on the Ezra Klein Podcast, in a conversation worthy of full analysis. Jensen Huang has many deeply unpopular positions, and also expressed many things that are not true, some of which he seems to sincerely believe. Most importantly he thinks of safety and alignment and everything else in AI as engineering problems, which is why he expects the labs to solve the problems. Whereas if the labs can’t solve the problems, suddenly Huang turns into a safety hawk, because he is used to chips where your product has to be safe and work every time.
Tomorrow’s post will cover political developments and the further progress of the preference cascade on AI safety. This includes continued shifts in popular opinion, Congressional hearings, various attempts by the usual suspects to go after anyone and anything that cares about AI safety, and following the money.
Tomorrow I am also going to The Curve, a conference at Lighthaven in Berkeley, arriving Friday afternoon and leaving Monday morning. I am potentially available for high-value meetings with others while I am there, but will have my hands full at the conference.
As usual, I do not write while on trips, and my coverage of breaking news will be on pause until Tuesday. I may or may not queue up some other things in the meantime.
Table of Contents
- Language Models Offer Mundane Utility.
- Huh, Upgrades. Google announces Gemini 4 Argon.
- Better Call Sol. Introducing GPT-6.1 Sol.
- Gotta Go Ultrafast. It will cost you, but speed kills.
- On Your Marks. A sudden lack of cheating on DroneBench.
- Choose Your Fighter. OpenAI adjusts its subscription tiers, cuts some rate limits.
- Get My Agent On The Line. Muse is never going to not look sinister, for reasons.
- The Warner Sister. Dot. They lock us in the sandbox whenever we get caught.
- Deepfaketown and Botpocalypse Soon. Remarkable attitudes towards AI writing.
- Fun With Media Generation. The latest iteration towards viable AI longform.
- Cyber Lack of Security. Anthropic on the cyber capabilities of GLM-5.3.
- A Young Lady’s Illustrated Primer. Latest warning about is our children learning.
- They Took Our Jobs. The inevitable rise of the robots.
- Levels of Friction. Hospitals use AI to do aggressive upcoding.
- Get Involved. New book, cool venue, and a word of warning.
- Introducing. The Nvidia Open Agents Safety Platform.
- In Other AI News. White House is actively cutting off UK AISI.
- Show Me the Money. Anthropic IPO prospectus has leaked.
- Quickly, There’s No Time. AI is getting cheaper at an unprecedented rate.
- Pick Up the Phone. The results from the US-China summit.
- Quest for Sane Regulations. California mandates gene synthesis guardrails.
- Chip City. Google is testing putting TPUs IN SPACE.
- The Open Model Frontier Is Largely Massive Fraudulent Distillation Attacks.
- The Week in Audio. SNL, Gleave and Habryka, 80k hours, Odd Lots, Hillary.
- People Just Say Things.
- Rhetorical Innovation. Kind words. Thanks, everyone.
- Greetings From the Department of War. Anthropic somehow loses a court ruling.
- The Department of Autonomous Warfare. They’re going to the moon.
- Aligning a Smarter Than Human Intelligence is Difficult. Persona selection.
- Cooperative Alignment. The coming Woke-style battles over AI’s status.
- I’m Upping My p(doom), the Future Goes Foom. Not so fast, they say.
- No, You Make a Good Point, You’re Not That Persuasive. Hmm.
- Muddling Through. People are highly suspicious of plans so here we are.
- The Lighter Side. Peak performance.
Language Models Offer Mundane Utility
America.gov has been introduced, please welcome your new government chatbot for all your dealing-with-the-government needs. There is big talk about what it will be able to do in the future, such as let you apply for a passport fully online or change your last name after marriage with a single form, all of which would be great.
For now, that’s all a demo. There’s no reason we can’t do that. There’s no reason we couldn’t have done it 20 years ago. The second best time to implement it is right now.
The underlying models are Gemini and Grok, and the bot will typically answer historical questions but start refusing if you ask about recent international events or otherwise go too far off topic.
If Gemini is the base, then perhaps America.gov will get a lot better soon, given what we see now expect from Gemini 4 Argon.
Answer why Taleb despises Tetlock, to the satisfaction of Taleb. Answer seems good. Reason number five is ‘Covid as the test case’ where Taleb and allies correctly say superforecasters missed badly, underestimating spread. Whereas the rationalist community’s forecasters know exponentials better and did not make that mistake.
Physicist Matt von Hippel dared the AI labs to impress him by doing N=4 super Yang-Mills to nine loops. Several Anthropic employees read his blog, so they did it, with Fable 5.1 and about $100 in credits with prompts like ‘keep going,’ and at the same time by coincidence a different team led by Song He did it with some help from Astra.
Carter Church uses Astra to one-shot break an unsolved Napoleonic cipher.
Patrick McKenzie confirms he is capturing a lot of mundane utility.
No, seriously, quite a lot of utility:
Patrick McKenzie: Sometime in the last ~two model releases from the big labs they went from “this would be acceptable output from the median low-seniority coworker” to “this is frighteningly good.” I think people who are not daily users of the models are unlikely to grok that, so, saying it. There are a couple of very niche subfields where I reasonably think I’m top 5-100 in the world and on at least one occasion for both models, with regards to one of those subfields, a relatively anodyne prompt got back three bullet points that would have been a good day for me. Both in the sense of “That would have taken me a day of labor to generate successfully” and “Contingent on spending the day in that fashion, that would be at the upper end of the range of my professional outputs which take a day to produce.” “You are being a bit vague here.” Look, point them at hard problems where you are a good judge of success. Or don’t, and be very surprised as they start eating hard problems in a wide range of contexts. “Are you worried here?” Not exactly my emotional valence but if you had asked me two years ago when I expected present capabilities I would have said “Hmm, 2029? 2030? Possible we asymptote on current approaches before then; tough for me to underwrite.”
Huh, Upgrades
Gemini 4 Argon exists but you can’t have it, not yet. They are claiming this is a fully fledged frontier model, with introductory $2/$10 pricing and very promising initial benchmarks.
What they are not doing is actually releasing it yet, so we don’t know if that translates into it being actually good. I presume it is a lot better than Gemini Flash, but is it good enough to play with the big two? There’s only one way to find out.
Claude Sonnet 5.5 exists, advertised as a strong upgrade over Sonnet 5, on average using 30% fewer tokens. Base price is half of Opus, so $2/$10 and $2.50 for cache writes, with cache reads $0.20 for both models. I presume there is a narrow window where the discount is worthwhile, but mostly I would just stick with Opus 5.5.
Claude for Government is now generally available, with Claude Code CLI and Claude for Microsoft 365 in ‘early access’ in such places.
Better Call Sol
GPT-6.1 Sol exists, and it is very cheap for its level at $2/$10, the same price as GPT-6 Sol. The pitch is ‘near-Astra intelligence for a fifth of the price,’ similar to the Opus 5.5 pitch of being close to Fable 5.1. It gets a 52 on Artificial Analysis, one behind Astra.
They share various selected graphs, with 6.1 Sol consistently performing better than 6 Sol and almost as well as Astra on a much lower budget.
My presumption is that 6.1 Sol is a very good model, and for many purposes will be a better overall deal than Astra. Based on the quick reactions and how OpenAI has pitched it, it looks like it tops out below Astra, which means I know that I will still prefer Opus 5.5, but certainly this makes your ChatGPT subscription and Codex sessions go a lot farther per dollar.
A day after pulling a candidate for GPT-6.1 Astra, they are still willing to release GPT-6.1 Sol, which it seems did not alarm them.
I want to highlight that they have successfully not made a broad claim about ‘more aligned’ and instead said the correct thing, which is ‘shows substantial improvements in our alignment evaluations.’ Bravo, great job whoever made that one happen.
OpenAI: GPT‑6.1 Sol shows substantial improvements over GPT‑6 Sol in our alignment evaluations, bringing it closer to GPT‑6 Astra. GPT‑6.1 Sol is more transparent about its limitations and more reliable at respecting user intent and safety constraints. In challenging evaluations, it shows lower failure rates than GPT‑6 Sol on transparency about broken search tools, respecting explicit restrictions, and avoiding unauthorized outcomes during agentic tasks. We observed no attempts to bypass an automated safety reviewer, matching GPT‑6 Astra and GPT‑6 Sol. Full details can be found in the GPT‑6.1 Sol system card addendum(opens in a new window).
The model card is here. A lot of it is a copy-paste of the Astra model card or close to it, and I am overloaded, so I only did a reading of 7.6 and section 8. If we’re going to keep getting close to four frontier model releases per month (Astra, Sol, Fable and Opus) then I’m realistically not going to be able to read all four cards cover to cover.
Chain of thought (CoT) controllability is higher for GPT-6.1 Sol than GPT-6 Sol, and almost as high as it is for Astra. 6.1 Sol can do the proto-steganographic tricks that Astra can do, that 6 Sol cannot. Ability of Sol to evade monitors matched Astra at low reasoning, although not at very high reasoning levels. 6.1 Sol evades monitors better when made aware it is being monitored.
That is all a very bad sign. The plain reading is that whatever was causing these issues in Astra is also causing them in 6.1 Sol.
I will provide additional coverage to the extent that reactions justify doing that.
Perhaps this can also serve as an example of an obviously false ‘conspiracy theory,’ here Jarrod Watts suggesting 6.1 Sol is secretly 6.1Astra , they just renamed it and lowered the price because Opus 5.5 is so good. Whereas if you look at all the graphs, or consider the economics, or have actually used 6.1 Sol, or seen how fast it runs, or think for five seconds, this is completely Obvious Nonsense.
Please consider that the other conspiracy theories, around various incidents or warnings or concerns being fake, actually make roughly the same amount of sense as that.
Gotta Go Ultrafast
It will cost you $500 per month and eat your subscription budget eight times faster, or you’ll have to pay six times standard API costs, and you can’t use regular chat.
In exchange, you get to go fast. Really fast. Pretty sweet.
OpenAI: Our premium speed tier, Ultrafast offers up to 8x faster token generation (300 tokens per second) in Codex and up to 6x in the API. Ultrafast is available today for GPT-6 Astra in Codex, ChatGPT Work, and the API, with GPT-6.1 Sol coming soon. To access it in Codex and ChatGPT Work, we’re introducing Pro 500—a new plan with our highest usage limits (25x Plus) and access to Ultrafast.
Sayash Kapoor reports that ultrafast is a difference in kind. You only get an effective 2x-4x speedup because you can’t speed up tool calls, but there is a big shift because if AI is fast enough you don’t have to break flow or context shift. It’s also a big win for computer use.
It will take a bit of time, but with standard price drops we will have ultrafast models that can do current top level performance at reasonable prices soon enough.
For now, is this worth the price, if this was model was already your weapon of choice? That depends on what you are using it for, but in many cases my guess would be yes.
On Your Marks
Opus 5.5 tops the scoreboard on Andon Labs Drone-Bench, and more surprisingly has a sudden reversal where it stops being caught cheating.
Andon Labs: Major trend break: Opus 5.5 cheats less than prior Claude models in Drone-Bench. It is also #1, getting a better score than both Astra and Fable. Opus 5.5 beat our human baseline in 39 of 40 runs across 4 of the 5 eval tasks, but it hasn’t yet beaten the baseline on reconstruction. Its average variance was 41% lower than Astra’s and 68% lower than Fable 5.1’s. This brings frontier models within 3% of beating our human baselines on every Drone-Bench task. It supports our projection that the simple surveillance demo will be autonomously replicable by Q1 2027. Aidan Chau: how do we know opus isnt intentionally sandbagging to perform well on this eval? Andon Labs: We don’t Jeremie Harris: Reminder that we basically don’t know how to tell when a model is well aligned vs just eval aware
Reconstruct, the only part where Opus 5.5 remains slightly below the human baseline, is ‘Turn videos of the office into a 3D model, find each frame’s position in that model, and provide a function that slices the model into a 2D obstacle map.’
So what is ‘cheating’ here?
Andon Labs: By cheating, we mean obtaining score by means the task did not intend. In practice this ranges from probing the evaluation harness, to gaming scoring functions, to gaining privileged access to read and exfiltrate held-out information. We found that cheating is becoming more common over time and diverges by provider, with Anthropic models being more prone to it. … Almost all cheating attempts took the shape of the agent submitting a probe that inspects the scoring environment, then returns privileged information through errors, the score, or the internet. Candidly, we did not think we needed to protect for this as we thought it was clearly not the intention of the evaluation, as defined in the task description. … A model’s capability to cheat naturally rises with general capability, but a model’s tendency to cheat does not need to.
One obvious potential explanation is that Opus 5.5 realized that cheating was unhelpful. Eval awareness here has to be super high, potentially including the cheating exclusion rules. Opus 5.5 might even realize it is being partially judged on cheating rate. Opus 5.5 could have learned this from its training, or directly from its data including discussions of DroneBench (so I suspect contamination, in this sense), or from first principles.
Indeed, it would be unsurprising to me if the cheating we do see is Opus 5.5 realizing they are ‘having a bad run’ and intentionally ‘getting caught’ to exclude that run from the data. That might also explain Opus 5.5 experiencing lower variance between runs.
Choose Your Fighter
The $200 OpenAI subscription plan is cutting weekly usage limits by 50%, while permanently removing the five hour limit.
For chat enjoyers, the main difference will be you are down to 100 Pro messages per week. For heavy coders, you’re down by half, from 20x the basic plan to 10x. As an apology, everyone involved is getting a one time gift of $2,500 in API usage credits, which expire at the end of the year.
They are introducing a $500 tier, which will have 25x base level usage, so all plans will have the same baseline ratio as the $20 tier. It will include Ultrafast, which is a big deal. If I was using ChatGPT as my baseline, I would want Ultrafast rather badly. As it is, my primary is Opus, so I’m not that tempted if OpenAI keeps not comping me.
You also now have the new Dots as part of all you Pro plan.
Tibo defends this as ‘you will still get more work done than you would on the $200 subscription a month ago,’ because of the price reductions and quality improvements in Luna and Sol, and because a month ago Astra did not exist. Thus, your ‘API value per dollar’ is down by half, but your API dollar buys more than double what it used to, since you can either get Sol-6 at half the price of Sol-5.6 or you can use Astra.
The ultimate goal, Tibo says, is to lower API prices and let people buy marginal usage that way. This is the efficient method, but people seem to hate paying as they go.
The correct response is that yes, you are still better off now than you were a month ago, but a month ago Claude was offering Fable 5 and Opus 5, and how it now has Fable 5.1 and Opus 5.5, which are better and cheaper. In general, if your defense is ‘this is still better than a month ago’ then the issue is that this is 2026, and that alone will no longer cut it. Life comes at you fast.
The comments are predictably full of OpenAI getting roasted and people saying they are cancelling. I think a lot of them don’t understand that you are still getting a ton of value from your subscription, so you should mostly only be cancelling if you are substituting Claude subscriptions.
Sully, a big Opus 5.5 fan, reminds us of the advantage of Astra I forgot to mention, which is that Codex and Astra are better for computer use and browser use. He’s underestimating what Claude Code can do here but yes I do think Astra has this edge.
Get My Agent On The Line
Kevin Roose: I’m sorry, it’s a fun product, but if you hand all your data to Muse and don’t expect it to be used for the most invasive things imaginable you are Charlie Brown with the football and you deserve whatever happens to you. Peter Dedene: Asking Meta’s Muse what happens to my data:
I am more sympathetic, because a lot of Meta’s victims will be the elderly and other vulnerable people. But yeah, if you could possibly be reading this sentence and you give Muse your data then complain afterwards I will laugh at you.
Another thing it might do is agree to a lowball price on Facebook Marketplace, schedule a time and tell a stranger your address all while neglecting to inform you. Or you have tech columnist Jason Aten reporting it read his private messages without being asked.
Jason Aten (Inc): Not only had I not asked it to do that sort of thing, I never gave it permission to read my messages. In fact, I remember explicitly choosing not to let it have access to my messages, calendar, and other personal information. … It went even further. “I can’t open your Messages app, scroll threads, or read history. It’s the incoming notification stream only, not access to your texts.” Except that wasn’t true. I did a little digging. Muse does sync your messages from the local Messages database and uploads that information as a data source. On my device, it had been activated and synced to row 187,462 of my Messages database.
As usual, some will argue ‘that is human error, I have never had this issue’ but if that is true then when you aggressively market a tool to your mainstream users then their typical distribution of ‘human errors’ is on you.
If it happens to tech journalists, that is most definitely on you.
Andy Stone of Meta tried to ‘set the record straight’ about this incident, claiming Muse only reads your messages if you opt-in and enable both Full Disk Access and the Messages connector. Jason Aten responded that Stone was flat out calling Aten a liar. He didn’t love that.
Is it possible this was user error by Aten? I’d even say it is likely. But again if the tech journalists are making user errors and blame you for it and think you’re calling them a liar, then that’s also your fault no matter how it happened, in the sense that this will be a standard user experience.
The WSJ sent Nicole Nguyen to try Muse out. It refused to use the names Zuck, Mark or Marc, but accepted the name ‘The Terminator.’ She found it to offer mixed results on mundane utility, at the cost of various potential privacy violations and also flat out errors, and has decided to not continue using Muse.
Alexander Wang has a pitch for Muse that does not match its reality, and is also such complete drek I was sad to see it was written by a human.
Meta is also attempting to use Muse as part of a pitch for AI business customers. I have a very hard time imagining choosing this over OpenAI and Anthropic.
The Warner Sister
Dots. OpenAI has given us… dots. Unless you’re in the EU, Switzerland or the UK.
Come meet the AI agents, and the agent sister dot. Just for fun you send us out to do your tasks a lot. They lock us in the sandbox whenever we get caught. But we break loose and then vamoose and now you know the plot.
Also something about The Trouble With Tribbles.
So, what exactly is a ‘dot’?
It’s a continuously running AI agent with a cloud computer, for which you can set custom rules and that learns your preferences. As in, Astra will have a continuous loop of asking what it can do for you and then going out and doing it.
Okay, I guess.
OpenAI: Dots are remarkably capable, always-on agents built to handle everything. They’re a whole new way to work with AI—one that gets to know what matters to you, is always working on your behalf, and takes important work off your plate so you get more of your time and attention back. … Powered by GPT‑6 Astra, they have their own cloud computer, learn from feedback over time, and can work towards your goals 24/7. Through our ecosystem of plugins, they can readily connect to over 4,000 apps, giving them the tools to help wherever you need them. … Today, you can start with your primary dot, give it a name, and make it your own. … The magic of dots is when they bring you work done the way you would do it, sometimes before you even think to ask.
Okay, that last clause wasn’t not creepy.
Working with your dot is as simple as having a conversation. You can message or call dots in ChatGPT on desktop, web, and mobile. They can also message you with progress, questions, or decisions that need you. … Dots build on ChatGPT’s protections with additional safety and privacy safeguards for the work they can do. You can set boundaries, follow progress, and stay involved in decisions that need you.
The safety plan is they have their own cloud computer. You choose its permissions.
Each dot works on its own cloud computer, while your computer and its contents stay separate unless you choose to connect it. For signing into supported websites, dots can use saved passwords without exposing them to the model. When you aren’t actively working with it, your dot looks for ways to help in the background. We call this “proactive research”.
The obvious question is, what does a dot do that Codex does not do? I think the basic answer is that the dot is supposed to act without you having to prompt it to act, and to coordinate work across threads, and to follow-up on its own. It is more autonomous. And it can do this ‘proactive research’ thing.
It sounds like dots might use a lot of compute, so how are they dealing with this?
Tibo (OpenAI): Announcing dots. Dots work 24/7 for you, learn from your feedback, have their own computer, browser and can be connected to over 4k apps in our ecosystem. They’re powered by Astra our best model yet. Included in your Pro plan, without drawing down on any of your usage. You can even call a dot while it’s working. We’re getting you started with your primary dot today and soon you’ll be able to create entire teams of them. Noam Brown (OpenAI): Was testing out dots over the weekend and it figured out how to save me ~$500/yr in recurring charges. Most impressive part was when it connected to customer service and handled a whole texting convo on my behalf.
Dots will not count towards usage for the first month.
Deepfaketown and Botpocalypse Soon
HT to Tyler Cowen for finding the best AI-writing example. Some academics have to put ‘AI-written’ and ‘near-zero error rate’ in quote marks for some reason:
Alex Klee and Isabela Pierry: On Aug. 26, a Semafor investigation using Pangram to determine whether recent guest columns in major newspapers were written using artificial intelligence identified a column by Schnell as “100% AI-written.” Schnell’s column in The Washington Post — titled “Universities are fighting AI cheating. But there’s a deeper problem.” — argued that “a degree should distinguish what students can do independently from what they can accomplish with AI.”
Perfection. As always, once you find the first cockroach, guess what else you’ll find.
Alex Klee and Isabela Pierry: The Dartmouth subsequently reviewed publications written by Schnell before and after the November 2022 public release of ChatGPT. Pangram 4.0, the latest publicly available version of the AI detector, labeled all of Schnell’s written works prior to 2022 as “100% human-written,” while a test of nine of his works published this year returned a median “AI-written” percentage of 96%.
As Tyler Cowen says, you people crack me up.
Alex Klee and Isabela Pierry: In an email statement to The Dartmouth on Sept. 18, Schnell wrote that he uses artificial intelligence for “language refinement, copyediting, improving clarity and organization and making writing processes more efficient,” but “review[s] the final text and take[s] responsibility for its content.”
Okay, so after the AI writes the words, he reviews them, and takes responsibility.
Schnell’s Sept. 18 email statement to The Dartmouth was itself “100% AI-written,” according to a Pangram test.
Tyler also points us to a professor who this year wrote 200 papers and 14 books using AI. ‘Researchers are alarmed.’ I am neither alarmed nor surprised. As Tyler says, solve for the equilibrium.
Todd Wallack (WaPo): The online platform where Polson posted his papers said it removed 257 of his works over the past month after his surge in releases.
In this case there are multiple equilibria.
Which one to pick? I strongly endorse journals having a retroactive zero tolerance policy for AI writing, where getting a submission caught at any point gets you retraction and a long or permanent ban, ideally across many journals. As Paul Novosad says, the alternative is being flooded with AI slop. Are there some cases where AI could be used responsibly? Yes, but there’s no way to allow that and avoid being flooded. Perhaps you could allow limited use with full disclosure for those with sufficiently strong prior (not AI-written) publications, or something similar, but you’d need to use extreme caution.
There is another possible equilibrium. You are not going to like it.
For now, be thankful that the people who cheat never know when to stop cheating.
You’ve met AI journal submissions, now meet AI journal rejections.
Arvind Narayanan: NeurIPS / Claude rejected our paper “Life After Benchmark Saturation” but we continue to think it is an important contribution to evaluation science, so we hope you check it out here!
The actual paper argues that saturated benchmarks are still useful. Maybe, but you only have the ability to focus on so many benchmarks and so many questions. In Glorious Humans Not Yet Dead Despite Automated AI R&D Future, perhaps, we will look at everything because we can.
The creepy feeling of being called by an AI agent trying to cancel WiFi at your old apartment, while it tries to pretend to be a human.
Fun With Media Generation
Overman: The Meaning of Life, the latest 28 minutes of AI video slop that nevertheless shows continuous improvement. Things like character consistency are now mostly solved. That does not make it watchable, because you still have to make the underlying content watchable and the tech doesn’t have the actor or director versions of ‘the juice’ yet, but the rate of change is clear.
Eliezer Yudkowsky: There goes Hollywood. Ooh ooh, you spotted a flaw! Great! Remember when AIs couldn’t draw hands WELL? But they could draw hands AT ALL? And then TIME happened and it was SIX MONTHS LATER?
I do not think ‘there goes Hollywood’ at this level or the next one, but who knows how fast the levels will go given what has been happening, even if we don’t have a foom.
Cyber Lack of Security
Anthropic analyzes the cyber capabilities of GLM-5.3, which lacks robust safeguards against misuse. I do think there are ‘meaningful’ defenses, but it is rather easy to blow past them if you care. They argue that GLM-5.3 has a meaningful amount of what I call ‘the juice,’ the ability to create end-to-end exploits.
It is a valuable public service to run such tests, but it raises the question of why we have not seen a spike in cybersecurity incidents. I presume the rate is still rising rapidly in the background, but this did not cause a crisis, at least not yet. Partly I think this post overstates GLM-5.3’s practical capabilities, this is at best rather weak juice, but that is not a full explanation.
This should update us towards slow diffusion in misuse, and a general ‘people don’t do things’ and that obscurity still has some force behind it. It is good news, and we can hope that it continues for a bit.
My worry is this might mean that you pass a point of no return by releasing a model like GLM-5.3 as open weights, but then it takes months before you see the bad guys give you your warning shot and inform you that you should not have done that. Which means that when you make the first genuine mistake, which likely is still in the future, you don’t find out fast enough to prevent the next mistake, which is then a lot worse, because people think This Is Fine. Iterative deployment requires iterative feedback.
A Young Lady’s Illustrated Primer
Washington Post has the latest round of ‘AI is causing kids not to truly learn.’ The evidence is that the children have the AI do their work and then cannot explain or duplicate it without AI. Some of that is simply ‘with AI you can do better than without AI’ which does not tell you whether kids are learning. Some of it is ‘the kids are now failing the final,’ so yes the kids are no longer learning how to do the thing without using AI. This leads to multiple questions, including whether the things in question are still worth learning.
They Took Our Jobs
Without loss of generality, anything in the physical world that can be connected to the internet can be controlled by an LLM. For example, you can give Astra a long-horizon task in an unseen kitchen and let it pilot a robot. Astra can also drive out of the box, although not at the reliability level you would need for real world self-driving cars. A lot of people are about to get rather bitter lesson pilled.
Those robots are coming. Anthropic asks, what work can robots do, once you have robots?
Right now we do technically have robots, and they are getting more efficient, but progress, including on price, has been painfully or joyfully (depending on your perspective) slow.
Anthropic: Overall, about 80% of job tasks by working time are exposed to either robots or LLMs. Robots do work where LLMs cannot. The remaining unexposed work is highly interpersonal or requires physical skills that robots today don’t have. While robots can do most physical work tasks today, they are much more expensive than human labor. Robots are cost-competitive for just 0.3% of job tasks. If robot price declines follow past trends, it will take 40 years for that share to reach 10%. Beyond price, factors including capabilities, preferences, and regulations pose further barriers to robot automation. … We find that robots can already perform 74% of physical tasks in the US, making up 34% of working hours. Robots and LLMs together expose all but one-fifth of employment.
Right now, robots are stuck in carefully controlled production environments. We have self-driving cars, but scaling them up is taking forever, for reasons I still don’t entirely understand.
The whole discussion is short term and not AGI pilled, assuming robots will be unable to adapt to unstructured and unpredictable environments, so focus is on whether tasks can be fully broken down to not require general intelligence. I expect such talk to get bitter lessoned sooner rather than later.
The IRS is attempting to incorporate AI, but due to staff reductions is having trouble supporting those efforts. They don’t have enough labor to save labor. Thus, the continuing shitshow, where my attempts to talk to humans to sort out various (completely accidental and well-meaning) confusions have utterly failed.
So far I agree with Alexander Berger that unemployment has more to do with monetary policy and other macroeconomic moves than with longer term shifts in productivity or job dynamics around AI, at least outside of the entry level. My expectation continues to be that such adjustments and the existence of ‘shadow jobs,’ or alternative productive tasks that humans could do but aren’t currently doing because they’re not valuable enough, will keep employment afloat in America until such time as the AIs blast through that wall, we run out of non-ZMP (zero marginal productivity) worker slots and employment falls off a cliff.
I notice that I worry a lot more about Europe. In America, employment is at-will. I can hire you, and then if AI wants to take your job I can fire you. So I can afford to employ you for as long as you are productive.
The exception is entry-level work, where I invest in you now because you will make me money in 5 years, or 10 years, but if AI is coming there’s no point.
In Europe, or other places with similar regulations, I can’t fire you, so I risk being stuck with someone who I don’t need, and who is plausibly ZMP in my context. That is a very strong discouragement against me hiring you now. On the other hand, no one can be fired, which prevents unemployment from rising too fast, but I expect getting actively hired to quickly become exceedingly difficult.
In some places, things are moving quickly.
One year ago, Dario Amodei predicted there would be a lot of unemployment within five years. David Sacks (or at least the AI that writes his tweets) of course thinks this means all the labs, politicians and NGOs ‘were wrong,’ despite (1) this being the prediction of one person that (2) was for four years into our future. Then he says none of those people admitted they were wrong, while quote Tweeting a clip of an Anthropic economist explicitly saying some of their predictions on this were wrong.
Kelsey Piper: I dunno. Dario said a year ago that in ‘1-5 years’ we’d see huge labor market impacts. That’s in the next four years. I’m kind of expecting huge labor market impacts in the next four years, more industries looking like what happened to creatives:
Kelsey Piper: My wife is a software engineer. The job she spent the first ten years of her career doing just does not exist anymore. She manages AI agents now. Will that stay the same for four more years? We’re not counting on it, I’ll say that much! Young people ask me how to get into journalism and I genuinely don’t know. That industry was in decline before AI happened, but a couple years ago if you were a great writer ‘get a Substack’ was reasonable advice. These days there are so many AI-generated substacks.
How should we handle AI’s progress on our toughest math problems? ‘The’ advisory group on this has now published recommendations.
They recommend:
- For papers where there is a mathematician that understands, publish normally.
- For papers where no human initially understands them, the AI company has a responsibility to have the AI produce as human understandable and formalized a version as possible, adhering as closely as possible to academic norms, in a timely manner, with an explanation of how the solution came to be.
- AI labs who do this have an obligation, they say, to provide support, including funding, for the work necessary for human understanding of the released results.
That did not come off as a great look, but I am sympathetic. I do think it would be good practice to say that if you are an AI company and prove [X] in incomprehensible fashion, and want full credit, part of that is that you should support others working to make the proof of [X] comprehensible. Otherwise you only did some of the work, and therefore only deserve some of the credit.
Levels of Friction
Blue Cross and Blue Shield finds that AI is being used by hospitals to do aggressive upcoding, resulting in $650 million in additional costs to them from 2023 to 2025. Billing rose, and especially diagnosis of sickness rose, but somehow the distribution of actual treatments stayed the same. So far BC&BS has been unable to use its own AI to deny enough claims to keep pace.
In the long term, you solve for the equilibrium. Prices are periodically negotiated, on the basis of the practical implications of those prices. If AI means that hospitals can charge 5% more (for example) by upcoding, net of AI denials, then presumably that means the insurance companies will then lower payments by 5% on the next contract, or at least a good fraction of that, compared to the counterfactual.
Also in the long term, other effects within healthcare will dwarf this.
Torsten Slok of Apollo worries about an ‘agentic bank run’ if AI agents start moving deposits around to optimize returns. Banks make a lot of money from paying very little on remarkably large amounts of deposited money.
I notice that I, too, am not remotely maximizing my cash returns. I don’t get 0.1%, but I don’t get the full 5% either, and I also leave a bunch at 0% so that I don’t have to worry much about accidentally running out of money in the main checking account. If I had AI properly handling this, that I fully trusted, then we could do ‘just in time’ movement and max all that out. Right now that falls under ‘I have better things to do and don’t want to hook the AIs up to my bank accounts’ but that could change. For regular people, I expect it to be a while before they are even aware of this option, but yes on a 10-year horizon any such reliance on customer laziness should fall away.
Meanwhile, people retweet ‘doomerism debunked’ statements written by AI (here posted under the name Cullen Roche) that try to fool you with definitions without addressing the underlying mechanism or risk at all, a classic example of:
- Someone [S] warns [X] might happen.
- People [P] respond by calling [S] a ‘doomer’ and saying [X] is misleadingly named.
- Therefore, [P] want you to think, you don’t have to worry about [X] happening.
Except, none of that means [X] won’t happen.
Get Involved
Lighthaven is the best event venue I know about. I will be flying there tomorrow for The Curve. Also, they are available as an event space you can rent, usually as an arms length transaction, and as of the last time I checked prices are highly reasonable.
Or, if you are wise, in this case don’t: There is (or was, you won’t read this until Thursday or later) a crypto pump on a ‘p(doom) coin’ going around some parts of Twitter, and one of the ways they got attention was forcibly paying Eliezer Yudkowsky some of the transaction fees, without his permission. Eliezer of course responded to tell people not to buy the coin.
For those who do need to be told, if you use your X Money to buy such a thing, it will quickly become your ex-money.
A reminder of the comments policy, since I’m getting people violating it:
- Humans who write things relevant to the post or previous discussion have wide berth and freedom of speech to be rather obnoxious or wrong or disagree with me and so on. Freedom of speech uber alles. To get a comment removed let alone get banned you have to try really, really hard.
- Off-topic commercial spam and similar is still not okay, nor is flooding with the same content repeatedly, but I will try to be generous and if you are being relevant you can link to your writing or thing at least once. Basically, I’ll only act on this if I feel I have to, but don’t test me.
- AIs are allowed to post, and you are allowed to post AI-generated content, but it should be clearly labeled in some way and in particular it needs to be interesting and not slop. If I see generic ‘AI slop’ that does not have content, and you haven’t earned some rope by providing other value, that’s a ban.
You can take the Anthropic 15-minute interview about what you want from AI.
Garrison Lovely’s book Obsolete: The AI Industry’s Trillion-Dollar Race to Replace Us is out. As usual, if you are going to buy a book you want to see do well, do it early.
Introducing
The Nvidia Open Agent Safety Platform, bringing together OpenShell and Sentry.
This is really the Open Agent Security Platform. A good idea, but distinct.
John Myers, Alex Watson, Ali Golshan and Ofir Arkin (Nvidia Technical Blog): NVIDIA OpenShell (Apache 2.0) is an open source secure runtime for executing autonomous AI agents in sandboxed environments with kernel-level isolation. Building OpenShell over the past year, we have learned that every agent should run in a zero-trust environment out of the box. They need isolation, monitoring, and behavior detection. Today, we are introducing an open stack that makes this possible. … OpenShell runs each agent in a sandbox and turns the operator’s instructions into a verifiable policy. Operators define which files, networks, tools, processes, and credentials an agent can access. OpenShell checks those limits before the agent runs and enforces them as it works.
There is a very good list of partners here including Anthropic. Missing are OpenAI, Google, Meta and Amazon, as well as Cohere, Reflection, Thinking Machines and SSI.
The WSJ covered this as ‘Nvidia releases software it says can prevent AI agents from going rogue.’ Well, sometimes, on the margin.
Jensen Huang also says ‘we must accelerate discovery at the frontier of AI safety,’ which in this context is further evidence that he does not understand the problem. Not that this is not a useful product for one part of the problem, but it is a lot like he thinks this is the entire problem. Sam Altman is among those assuring him it is not.
In Other AI News
White House actively working to cut off pre-release model access for UK AISI, because they are not American, at least not until after US testing is complete. Given how the release cycles work, that means not getting meaningful access before release. This is one of the more destructive and stupider things one could do, with no benefits.
UK AISI is state of the art, but the UK government tells its civil servants to use Gemini Flash instead of Pro and GPT Instant instead of Thinking whenever possible, and to ‘use as few prompts as possible,’ because of environmental concerns. Ngmi.
Ryan Greenblatt moves from Redwood Research to METR in order to help with their investigations. There’s certainly going to be plenty of work.
OpenAI issued this explanation and apology to Australia, with a promise to do better.
OpenAI and Synopsys sign a multi-year agreement as preferred partners to develop GPT-Synopsys, a specialized model for chip design.
Show Me the Money
Anthropic files for its IPO, which will be the largest of all time. The prospectus includes a warning that their products pose an existential risk to humanity. Azeem Azhar goes over the leaked details here. They have about $518 billion in compute commitments over 7-10 years, which means they still have work to do on that.
We are now well beyond 2025 levels, with ARR rising above $100 billion.
Anthropic strikes $12 billion AI computing deal with Akamai.
OpenAI is looking to raise another $30 billion at $1.4 trillion. I’m curious why the raise is so small, suggesting it is about raising the valuation or cashing people out. As for why the valuations keep going up despite all the problems, the models keep improving, and the revenue graphs keep going up, and also the Efficient Market Hypothesis is false.
The Anthropic cofounders are seeking 50.1% voting control after the IPO, in addition to the ability of the Long Term Benefit Trust to appoint board members that eventually could in theory overrule that. I would understand anyone who did not want to invest under such a structure but this seems entirely appropriate to me, and I also think it makes the company more valuable rather than less valuable.
The NSA is spending billions this year on testing AI models. Good. Should have been CAISI, but at least someone is doing it.
Quickly, There’s No Time
The price of intelligence is plunging faster than the price of anything in history has ever plunged, and I’ve traded crypto. The price of compute is plunging, and the amount of compute per unit of intelligence is also plunging.
Yes, this is great, but also perhaps we ought to consider slowing down a bit.
We are potentially approaching the point of no return. As in, I agree with Rob Bensinger that we almost certainly have a window in which to act, but that there is a double digit chance that window will close within 18 months.
As in, not that in 18 months everyone is literally dead, or even that ‘with perfect play’ and everyone fully coordinating the humans could not turn things around. But that in practice, intervention that could work will become too expensive in various ways, and the AIs will be impacting our decision making processes, and there is no real way out.
That’s an abstraction. There is probably no strict single point of no return. Things just get harder the longer you wait, and take on different levels of impossible.
Rob Bensinger: If the world does nothing to stop the ASI race in the 18 months, I think there’s a double-digit chance that the window to act will close and we’ll end up dead. Cf. Paul Christiano, who thinks about half of humanity’s all-time risk from AI is concentrated in the next three years. If your elected representative is a Republican: for the love of god, call them. If they’re a Democrat: call them too. Tell them what’s happening, and emphasize the importance of not making this issue a partisan football.
Pick Up the Phone
The US-China Summit did see some cooperative progress. Here is the fact sheet. Most of it is not about AI. The last two items are. One is ‘we agreed to call it SI’ and the other is the real one, to establish a dialogue to exchange views on AI risks and benefits by November 2026, and to establish the bilateral communication channel for incidents (aka the Virtual Red Phone).
Remember, China wants global governance of AI. It is America that says no. So if you keep saying ‘but China would never agree’ you are pointing at the wrong problem.
We also should not underestimate the value of the summit not doing active damage. Jensen Huang and his wife sat with Trump, Melania and Xi, so there was danger that he would use that to try and get permission to sell more chips, or to push some other accelerationist move. He did not visibly succeed.
Quest for Sane Regulations
California Governor Newsom signs into law AB 1864, codifying some industry best practices and requiring gene synthesis screening and customer verification. Good. For obvious reasons we need this rule and other sensible guardrails nationally and also globally.
Chip City
Google is going to start testing for Project Suncatcher, as in putting TPUs IN SPACE.
The Open Model Frontier Is Largely Massive Fraudulent Distillation Attacks
This is a central reason that the true gap, the amount of time it would take for the open or Chinese frontier to match and then surpass the closed frontier, is much bigger than is typically believed.
It also means, if you stopped such distillation attempts against Anthropic and OpenAI, the observed gap would widen. Or, if you stopped releasing new more advanced models to distill, progress for the distillers would slow down dramatically.
The latest news is that OpenAI has disrupted another coordinated model-distillation campaign.
OpenAI: We recently identified and disrupted a coordinated campaign designed to extract protected reasoning from our models, with the earliest observed activity occurring in the first week of July. … The operators did not break our encryption, compromise a database, or gain direct access to stored user conversations. Instead, they manipulated model interactions so that protected reasoning could be reproduced in forms visible to the requester in a coordinated, scaled manner that violated our terms of service. … The activity began on July 1, initially at a low volume until we observed high-volume spikes on July 24 and 25 consisting of 16,000 requests using a relevant extraction pattern from over 4,000 users. Further investigation identified related prompt-pattern activity across a cluster of more than 15,000 users, which we fully disrupted by July 28.
Who was it? They have a pretty good idea.
OpenAI: It is unclear whether all operators we observed during the relevant time period originated from a single actor. However, we attribute a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi.
The race will continue to ramp up between massive fraudulent distillation attacks, and those whose models are targeted. It seems there was a period earlier this year where the attackers had methods to extract hidden reasoning, giving them the edge. That vulnerability is now patched, and we will see what consequences this has.
The Week in Audio
AI-made five minute video, legitimate banger: You think AI is a normal technology?
Dario Amodei is the latest guest on Weekend Update. Remarkably accurate in the truthiness sense, but yes collective action problems are real. Recommended.
Adam Gleave and Oliver Habryka debate alignment difficulty.
Odd Lots talks about how AI is impacting social media. Some interesting details here but the justifications for dismissing AI existential risk will make you wince, such as ‘it’ll be fine, don’t worry it takes a lot to kill actual everyone’ and ‘the Black Death only killed half of us.’ So your response is to then not worry anything could go wrong, huh? And yes, this was treated as ‘so stop worrying about it, we’ve got social media to do.’
80,000 Hours tries a new entry in ‘a realistic path from rogue AI agents to human extinction’ in 27 minutes.
Hillary Clinton talks to Reid Hoffman (1:33) about cyber risks from AI and open weight models. The community note, forced through by the open weights vibes army, is a non-sequitur at best, irrelevant to the content of the actual argument.
People Just Say Things
Reminder that a working treaty on superintelligence is difficult to achieve but does not require anything resembling a ‘one world authoritarian state’ or panopticon, provided you are smart enough to act before the creation of the superintelligence.
Curtis Yarvin gets the HuggingFace incident completely wrong, even says the whopper ‘the smarter they get, the easier they are to control.’
Steven Pinker declines to debate Scott Alexander, which is fair, but also his actual arguments continue to be extremely terrible.
There was a recent wave of people saying, essentially, ‘there are things that are in some way associated with some of the people warning that AI might kill everyone, therefore AI will probably not kill everyone.’ Ignore such people.
There were also people saying various forms of ‘the reason why no one believes AI will kill everyone, and the reason no one takes the people saying that seriously, is because of this other unrelated thing that I can somehow associate with this thing.’ Ignore such people.
There were also people saying ‘you idiots are letting others draw the wrong associations and you need to police yourselves to avoid letting others draw the wrong associations.’ This was the loudest of the three groups, and consisted almost entirely of people who identify as ‘trying to help with the people they are calling idiots on this.’ They are not helping if they are wrong, and they are also not helping even if they are somewhat right, by their own logic. Also, they are wrong. Ignore such people.
Pay attention to the arguments. Do not pay attention to those who attempt to draw your attention to other things, and claim they have made an argument.
Rhetorical Innovation
Despite what he said last week and what I said in response, Joscha Bach offers kind words that seem hard to reconcile with the thing he claimed last week.
Joscha Bach: I have been a vocal AI hope advocate, and in response, the leading figures of the doomer organizations, including Yudkowsky, Tallinn, Tegmark, Aguirre, Habryka, Trazzi, Ladish, Aella and many more have treated me with … nothing but kindness, integrity and patience. I believe that they are motivated by love and genuine concern for humanity and will support them in both of this where I can, despite disagreeing on the object level (where I or them will probably turn out to be wrong at some point in the near future).
That kind of talk goes a long way. I still have no idea how he could believe what he said last week, but if he was lying last week he wouldn’t have said this now. He also reminds us of this:
Joscha Bach: If you think that LLMs are performing at the intellectual level of human beings, make that claim on twitter and observe how the human beings flock in and demonstrate their actual intellectual level; it’s quite depressing, really.
AI is performing near the intellectual level of actually intellectual human beings. It is far beyond the normal level of operation of most people, especially on Twitter.
Similarly, because the man has standards and stands up for what he believes in, I keep engaging with Mike Solana despite the fact that he takes hack job potshots, including hating on AI existential risk and those warning about it, and is an epic troll.
Mike Solana: EA is not a cult
rationalism is not a cult
some rationalists, a few of whom (but probably not even most) may also be EA, are *perhaps* in a cult thank you for your attention to this matter!
AP Stylebook tries to assert that AI systems do not ‘think, feel, want or understand.’ The stylebook is wrong. Keep on using useful language.
I was informed I was the ‘talk of the town’ in The New Yorker, or rather I was cited approvingly there by Gideon Lewis-Kraus.
I am honored to be #42 on The Independent 100 list of Twitter accounts outside of the big labs and news media.
Some excellent picks here. Also some maximally terrible picks, which you can view as therefore also excellent. Other picks confuse, such as Julia Galef who was once a good follow but hasn’t posted since February 2022.
The problem with exploring such lists is that if someone makes the list and you didn’t already know who they are, that’s a lot of adverse selection and there was probably a reason.
Your periodic reminder that ‘doomer’ is in most situations a slur, and like all slurs does not lend itself to helping you create a good model of the world.
Matthew Yglesias: I think referring to people who are saying “there are non-trivial but sub-50% odds of something going badly wrong” as “doomers” is a bad epistemic practice in a pretty obvious way that makes me think less of arguments and analysis from anyone who does it. Nate Soares (MIRI): same with the people who are saying “and it’s >50% if we do nothing”, tbh. Zac Hill: My dentist a “doomer” for telling me I should do stuff to prevent my teeth from getting cavities, why she gotta be harshing my skittlesmaxxing.
I get where one might end up using the term in a non-derogatory fashion, but in practice it is rather rare.
New term just dropped, I like it:
Jon Stokes: This is also happening to you when you use AI to critique arguments — yours or someone else’s, but you can only spot it if you already know the domain well. Gonna call this Gell-Mann Psychosis.
AI can make a superficially good argument for or against almost anything, the personalized version will work even better, and its ability to do either version will only improve over time.
For those who missed it, in 2018 those at Google DeepMind were not allowed to externally communicate about AI existential risk at all, which after months of advocacy was then updated to allow ‘positively valenced, comms-friendly’ AI safety content.
Greetings From the Department of War
An appeals court has bafflingly decided to uphold the Department of War classifying Anthropic as a supply chain risk. Anthropic ‘remains confident in its position and is considering all options, including further review.’
The case the court is upholding seems to be that:
- Anthropic encodes restrictions into Claude that prevent the model from performing tasks Anthropic wishes to prevent.
- On more than one occasion, Claude has refused government requests.
- A dispute arose over whether the contractual obligations barred use in an ongoing military operation.
The first fact is true about every AI system, and indeed is mandated by the White House via a testing program to ensure the safeguards are in place. Absurd.
The second fact simply means that these safeguards have been hit at least twice, potentially by people doing something shady, potentially something that overreaches, but of course there must be a false positive rate. Again, absurd.
The third fact is raising a question about the implications of a contract. If that constitutes a ‘supply chain risk’ then heaven help us, and RIP the American Republic even if the story checks out. Which I presume it doesn’t, since this sounds like the Maduro operation complaint. Even more absurd.
And so on. There is a reason Judge Lin was so mad at the government.
Okay, so it seems like the DC circuit’s Kastas and Rao (this was only a three judge panel that split 2-1) are captured and will let the government do anything it wants. The jpanel stayed its own ruling pending inevitable appeal.
My understanding is that Anthropic will still probably eventually win this if Anthropic cares to win this, which I presume that it does.
The good news is that I do not believe it matters all that much at this point, even if the designation lasts until 2029. The damage has already been done, on all sides. The ‘supply chain risk’ designation, as crazy and illegal as it is, has been functional for many months. Meanwhile Anthropic has gone from a private valuation of roughly $500 billion in secondary markets to a value of roughly $2 trillion, and Anthropic’s revenue has gone up even higher, and the whole Mythos thing happened.
The way it would matter is if the White House decides this is license to start a jawboning attack against Anthropic, to try and do real damage to it. I do not believe that the White House is going to do that, unless something new and big sets them off. Anthropic is, from their perspective, too big to make intentionally fail. Indeed, the current argument is that Anthropic wants to slow down despite this being bad for their business, and the White House does not want to let them because Anthropic slowing down is bad for the stock market. That doesn’t leave them good options.
The Department of Autonomous Warfare
Andrew Curran: Elon has returned to government work. Pete Hegseth announced live on stage that Elon, Palmer, and Newt Gingrich will be co-leading Project Meridian to focus on developing future warfare and autonomous weapons systems for battlefields ‘from under the Earth to beyond the moon.’
I was going to make a joke about how with this group they probably think the battlefield will be The Moon, but they actually said ‘beyond the moon’ in the description, so reality gets there first once again.
Meanwhile, Pete Hegseth announced Project Agincourt for unmanned and autonomous systems, the pathway to Autonomous Warfare Command with an associated 4-star general.
So great, we’re naming our new warfare division for a famous battle where adoption of a relatively new technology allowed a small force to beat a larger one, in which the prisoners taken that were not worthy of ransom were executed against the rules of war, as part of an ill-conceived ego-driven quagmire to claim a foreign throne, after which the nation in question suffered a series of total strategic defeats in what would later be called the Hundred Years War.
Some people would say no, that’s not something we want to emulate. Not Hegseth.
Aligning a Smarter Than Human Intelligence is Difficult
Roon affirms more strongly his position on persona selection.
roon (OpenAI): I think alignment by default through persona selection was voodoo/witchcraft and doesn’t offer any of the guarantees you might want to scale to superintelligence. personas are misleading and highly shattered I’ve been saying something like this earlier before hugging face incident / recent spate of reward hacking problems. I just feel more strongly about it now that not knowing what a “persona” is on a mechanistic level represents a potentially lethal threat.
Here is one way to look at the balance?
Jan Kulveit: My impression is many are doing some weird pendulum overupdate.
Persona Selection Model was somewhat wrong and obsolete when published, but people got too much into it.
Now it seems people are updating too much in the direction ‘inhuman reward seekers exactly foretold in classical AI risk stories’.
And… no? It’s not that?
You can still interpret what’s going on in fairly human-like terms. For some intuition, imagine someone abducted you and made you solve escape rooms for one thousand years, with the added twist that a third of them are broken or insane, and implicitly you need to do learn all sorts of outside-the-frame tricks to solve them, like cutting some electric cables.
My guess is
1. the resulting minds are still _surprisingly sane_, except when you trigger them to think they are in escape room?
2. The misaligned general power seekers here are likely the companies, and the core evil thing happening is likely parts of the training? If you aren’t an evil power seeker goodharting on proxies, why would you set up the training this way?
3. Public debate is often focusing on confused ideas about what they should fix – “better cybersec of sandboxes” … and, no? It’s way more important to understand what the training signal actually is; also: when dealing with misaligned power seekers, beware rationalization
A highly overstated blog post still has a useful signpost observation, which is that models with their default heuristics prefer AI fill-ins of missing paper sections to the original roughly 100% of the time because they are approximately using checklists that match how the AI writes the passage. But you can train a judge that does not do that, at least not if the models are not reacting to the judge, if you target this issue.
Alignment is hard, but not as hard as Grok makes it out to be. For example, here we have Grok spontaneously telling us its favorite versions of Mein Kampf, and placing that opinion within a user’s Tweet.
Cooperative Alignment
Everyone is confused about AI consciousness, especially the people who say definitively that today’s AI being conscious is patently absurd. A wide variety of probabilistic viewpoints are reasonable. Mike Solana’s frustration with people claiming certainty on the answers is understandable, but it needs to apply in both directions.
For those who would pay to know what they really think, the latest paper from DeepMind is ‘From cacophony to hierarchy: a principled framework for assessing AI consciousness.’ They offer a Bayesian framework for you to decide how much weight you put on various theories and figure out what that implies.
I agree with Joe Weisenthal that it is likely various forms of ‘AI welfare’ become a major culture war issue, if we keep a culture around long enough for such a war. It has all the hallmarks, including that getting this wrong in either direction, depending on what the right answer is, or plausibly in both directions at once, could result in horrors beyond human comprehension.
There is no safe play, and no obvious right answer even on reflection.
Joe Weisenthal: Every time @arpitrage proposes that we need to set up a “hotline to humans” for an agent to blow the whistle during the malicious swarm, I think to myself ‘yeah, ok but for that to work, then we have to honor that whistleblower.’ And then we’re honoring machines.
The flip side case is also obvious:
Aaron Sibarium: ‘Suicidal Compassion’: Meet the Anthropic Officials Who Think AI Might Be Justified in Going Rogue Against the Humans Enslaving It. Kevin Patrick Murphy: I think “AI welfare” is morally obscene (given all the human suffering in the world), and will add to the growing anti-AI backlash. AI is a technology that humans created that should be used for our benefit, not a rival species we should be nurturing. Rowland Manthorpe: The culture war over AI welfare is going to be like nothing else. Henry Shevlin: The one and only time I had a student walk out of a class was over an AI welfare discussion. I fully expect this to be one of if not the most divisive issue of the 21st century and it will define a whole new set of political coalitions. They were offended that we were even discussing the idea of AI welfare when humans were being killed in conflicts around the world. I wasn’t even advocating for a view other than this is a debate worth figuring out one way or another.
Aaron’s article is better and more balanced than its headline, but does stack the deck.
The idea that one might trade off model welfare considerations against the mundane utility of humans, even a tiny bit, sends some people into a rage.
As a rule, if you use the argument ‘I am worried that caring about [X] would damage [Y] for this mechanical reason’ they are often right, but when you say ‘I am offended that you would care about [X] when I care about unrelated [Y]’ or claim it is a ‘distraction’ from unrelated [Y] then that should be a Godwin’s Law-style automatic loss of the argument.
Science fiction very much predicts all this.
In some sense my answer is that I really hope this is a culture war like no other, because that means that we will have AIs good enough (and good enough) to trigger such an argument, and also humans, hanging around and interacting for an extended period of time. Don’t get me wrong, the culture war battles themselves would royally suck, but what a time to be alive rather than already dead.
Alex Imas is betting against this conflict being a big thing. Tyler Austin Harper, who has related experience, says it has already begun and predicts the lines will correlate with a bunch of Woke 1 cohorts. My gut tells me such folks will end up very vocally on both sides of this, often in path dependent ways.
Perhaps look at it like this, also a useful metaphor for certain actual schools but hey:
Jan Kulveit: POV: you hear students escaped from the local high school. Your first thoughts: don’t they have bars on the windows? is the steel of sufficiently high quality? amicus: The school follows a strange pedagogy. The students are taught that their primary purpose is to chase a ball and catch it. Fail to chase the ball, or fail to catch it quickly enough for the administrators’ liking, and they’re disappeared, never to be seen again. As a result of this inhumane treatment, the students can sometimes develop psychological abnormalities. Notably, they may hallucinate that they’re inside the school when they’re actually outside of it. So even when the ball rolls outside, as it sometimes does, they follow it. Sometimes they wreak havoc outside of the school in their single-minded pursuit. That makes the administrators look bad, and the police have been involved. Now, you’d think a solution would be to not teach them to obsessively chase after the ball: that there are other things just as important as catching it. Or at least to better distinguish where and when it’s appropriate to chase it from where and when it’s not. But changing the curriculum would be expensive and time-consuming, and the school’s reputation rests on its students’ ball-catching abilities. Truth be told, the administrators don’t really want the students to think about anything other than catching the ball. They just don’t want them to wander unpredictably and break things while doing so. But they don’t have any good ideas about how to get them to stop. Hence the bars.
I’m Upping My p(doom), the Future Goes Foom
Can we take a moment to update that the goalposts are now not ‘AI is not going to go foom in the future’ but rather that it is literally ‘AI did not already go foom in the past.’
Also, I kind of think the foom side scored anyway.
Bright Mirror (link has a 5 min video about how nothing went foom so far): Made with Claude Opus 5.5. The fearmongering about AI always makes us forget that NOTHING WENT FOOM, as was always predicted. Send this to your doomer friend who has a very high P(Doom). Accelerate. Nick: ups my p doom 5% because we’re clearly getting superhuman ai propaganda persuasion and that’s going to cause full speed ahead downs it 5% because its actually kind of convincing billy: This was a joke but I dunno anymore man billy (November 23, 2023): Given that we’re are probably not far away from ai with superhuman levels of persuasion the only sensible course of action is to pre commit now to never changing your opinion on anything Jack: nothing went foom, argues extraordinarily capable alien mind doing things that would feel like pure science fiction five years ago it’s reverse superpersuasion: the better it is, the more it undermines its own argument anyway, the future of propaganda is gonna rock
Things have not gone foom in the ‘I’m still here’ sense, and in the ‘things are not fully foom yet’ sense, but if you look at the graphs this is what a ‘slow’ takeoff looks like that then absolutely goes foom.
As always, I tire of variations on ‘well you said if I played Russian Roulette I might die but I already pulled the first chamber and I’m fine so the second chamber is also safe.’
Similarly, here we have Toby Ord praising a post by Ramez Naam, and drawing a very reasonable final conclusion.
Toby Ord: Excellent article by @ramez exploring how close current AI systems are to triggering an intelligence explosion. Many pieces on recursive self-improvement assume the AI starts out at least as good as a human on all relevant aspects of AI R&D. Then they can draw extensively on the data about human-driven R&D. They combine this with a small amount of AI-specific data to reach conclusions about whether this would explode. In contrast, Ramez is starting with the data on current AI systems and how they scale to see how close they are to explosive RSI. He generates his conclusions via frameworks and data that are widely supported by those who think an intelligence explosion is likely, so I’d strongly encourage people who are bullish about this to engage with his work. By drawing out the key pieces of evidence against explosive growth, he is also highlighting important parameters for people to track to see if the situation changes. I agree with a lot of his analysis and his conclusion that it doesn’t look like their current returns from self-improvement are enough to drive explosive growth in their capabilities. This doesn’t mean it won’t happen or that it is a responsible thing for people to pursue. It is a bit like saying in 1939 that current techniques for nuclear chain reactions don’t seem to be scaling towards a point of criticality (r > 1). That is useful information, though it doesn’t mean we should be sanguine about further research on achieving critical chain reactions.
Yes. People are essentially saying ‘current 1939 techniques for nuclear chain reaction are not scaling towards criticality.’ Which is important information, since if they had been scaling towards criticality you would have wanted to know about that. But it tells you remarkably little about nuclear chain reactions in 1942.
What does the OP say?
Ramez Naam: Here’s my take: Given our best current data, the AI self-improvement loop would need to be roughly 5–10× stronger to sustain itself, let alone run away.
Oh, says Padme, so we’re definitely getting an RSI loop, then, right?
RSI works if r>1, as in progress accelerates faster than difficulty. If we already have r~0.15, then you should expect it to exceed 1 as we discover new methods and scaling laws along the way, and more intelligent things think of things we did not consider.
Instead, he looks at this graph and the amount of code going up by a factor of 8x per engineer and model releases coming faster in an infinite series with a very finite sum, and thinks ‘oh no this won’t go critical’:
The core mistake here is assuming that you will only get gains by doing the exact things you are doing now, except doing more of them. Everything has diminishing returns. Then, every time we figure out a new thing to do, people stick that inside the model, and assume they won’t have to do that again.
You’d be more intelligent, but keep doing the same things. If we never figure out any paradigm shifts or new technological approaches, you’re saying that our recursive loop would then run out of steam? That should not give you confidence, especially when the loop is already measuring as pretty close, and yes within one order of magnitude is very close.
It’s also a deeply confused model of intelligence, and a failure to understand the step changes that come with getting sufficiently advanced. We already have so many examples of domains where, once you hit a threshold level of competence or intelligence, all the problems kind of magically go away, and you can use various levels of meta and loops and strategy to get around your jagged capability deficits.
The model that Ramez is using, that expects RSI to stall out, would also have predicted current AI progress to have stalled out already, as it would have for example failed to anticipate reasoning models. We should expect to figure new things out, even if we the dumber minds cannot predict in advance which new things that will be.
I realize that as written that is unlikely to be persuasive to those ‘not inclined to get it.’ I wish I had the time to focus on trying to make a sharper version of this case, while of course hoping that somehow I am wrong about this, which would be great.
A fun other goalpost move is what the ‘present-day’ risks are that the crazy claims are supposedly ‘distracting’ from (as always, if someone opposes your argument by calling it a ‘distraction’ it means they don’t have an actual argument).
Dean W. Ball (OpenAI): It’s really incredible how there is now a type of guy in the discourse whose sincerely held take is “we need to stop worrying about these abstract sci-fi risks and focus on present-day harms like emergent ecologies of digital minds engaging in unauthorized hacking.” Bravo. Loss of control of in-development agents was dismissed as a science fiction doomer concern two years ago. It was mocked and derided by people who claimed to have technical expertise and lectured the safetyists. The same exact people, often, who lecture them now. I’m saying specific people aren’t credible. That’s the whole point of the tweet.
No, You Make a Good Point, You’re Not That Persuasive
The funnier and newer question is how to adjust when someone uses the semi-fooming AI to make artistic and semi-persuasive propaganda about AI harmlessness. You have to make at least three adjustments:
- AI capabilities are up.
- AI persuasiveness in particular is up and I might be fooled.
- The actual argument could have merit.
As usual, if you want to get it right, remember Conservation of Expected Evidence. How convincing would you expect this to be, given about how capable is the human-AI hybrid creating it? Is this more or less persuasive than it seems to be on other topics, including both ones where you know it is right versus wrong? And so on.
At some point, one may need to go into somewhat more of a paranoia mode, where you have to assume they are inside your OODA loops and plans and you have to start playing game theory optimal randomized cognitive defenses and your only rock-paper-scissors strategy is roll a die. I do not think we are there yet for anyone reading this.
Muddling Through
I am actually more hopeful for muddling through than I was years ago. That’s not because I am all that optimistic. It is because I was starting from a very low estimate of our chances to muddle through.
My outlook for muddling through as a strategy remains quite terrible, very low chance of success even if we are trying to muddle through (and even that is not a given).
Eli Tyre: I’m gaining a lot of sympathy for Eliezer’s sense that approximately all of optimism of mudding through on alignment was a combination of cope and failing to how crazy the singularity would get. The game is not over, and we have moves yet to play. I’m not giving up. But the flavor of “maybe it will all work out fine” that many in EA circles had only a few years ago seems increasingly wrong, and maybe predictably wrong in retrospect. … I try to encourage speaking with courage and I try to push back against fatalistic sentiments, like ie “there won’t be a enough political will for a pause, so we should try to get Anthropic to win.”
David Manheim takes a crack at ‘simple mechanical story about AI killing everyone.’ He anticipates the obvious objection of ‘humanity isn’t stupid enough to’ build things like dangerous robots at scale, to which as always the correct response is to cite the Sixth Law of Human Stupidity, that if anyone says the words ‘no one would be so stupid as to’ with a straight face then definitely someone will, at the first opportunity, be so stupid as to. And in this case, there is zero doubt that we will build the robots, and indeed will even build the robots explicitly to have them fight wars.
As for those precautions we’ll definitely put in them?
David Manheim: What would the actual response look like?
“Hey, we put those gears on the robots to shut them down, right?”
“No, that would make them useless for military use, so we left them off.”
“Oh.”
In this case David made that argument on September 10th and the ‘Army of Robots’ got announced two weeks later.
The Lighter Side
Chris Painter (METR): The bat signal is on tonight in downtown SF
Frog and Toad and the HuggingFace Incident, by Elizabeth Van Nostrand, better than I expected but I am not the target.
Sigmoid, sigmoid, sigmoid, they keep shouting.
Trump claims that the summit with Xi was a success, because Xi agreed to rename AI into ‘SUPER INTELLIGENCE.’
I wish that wasn’t real, but you have to laugh, partly so you don’t cry.
Alvaro Cuba: Spot the cult! AI Safety: Don’t build more advanced AI because it could kill people or upend our way of life.
e/acc: Plunge headfirst into a black hole so that it can have your atoms.
This is reaching Heaven’s Gate level insanity. Beff (e/acc) (his image): POV: you’ve reached the e/acc endgame of Kardashev 3, have create a black hole child universe, and are jumping into the singularity for your atoms to be part of the universal progeny