Is it true that “AI Chooses Nuclear Option in 95% of War Simulations”?

In a message entitled, “AI Hype Goes Nuclear: Debunking a Headline-Making Preprint,” Geoff Holtzman writes:

On the last day of February, I [Holtzman] met some friends for dinner. Earlier that day, the U.S. had struck Iran with the help of Anthropic LLMs, initiating the ongoing war in Iran. The previous day, the Department of Defense had cancelled a $200m contract with Anthropic, signing a less-constrained contract with OpenAI more-or-less simultaneously. And according to my friend, researchers had just found that AI used nuclear weapons in 95% of war simulations. Wait, what? At first, this last news flash struck me as irrelevant, probably untrue, and not worth writing about. But then, in the ten days after New Scientist first reported on the preprint, Axios, Newsweek, and many other name-brand outlets ran the story as headline news. Newsweek’s headline was the most concise: “AI Chooses Nuclear Option in 95% of War Simulations” In the weeks that followed, largely unrelated pieces in The New York Times, NPR’s On the Media, and Vox (twice) have offhandedly mentioned this “95%” stat in a single sentence, as though it’s not even worth questioning—which is exactly how pseudoscience becomes ‘common knowledge.’ At the very least, the statistic seems to reflect popular sentiment. In one Reddit thread, the New Scientist headline has 2,200 upvotes—though that doesn’t mean much, since the same headline garnered 36,000 upvotes in another. Even Tom’s Hardware scored a respectable 4.9k with this story. And why does any of that matter? Because these stories encourage readers to ignore all the clear and present dangers posed by the LLM industry. In fact, the New York Times piece invoking the AI-nukes study is credulously titled “Data Centers Are a Distraction.” Besides drawing attention away from more important matters, this narrative serves as a clever sort of hype for LLMs’ ostensible intellectual prowess. The way the story has been covered tends to imply that the amoral, non-emotional nature of LLMs is important because LLMs are so intellectually powerful—or at least intellectually real. I could go on about the causes and effects of this news pollution (as I did here), but for now, I just want to reveal how teachably-bad this study was at almost every level. Bad design, bad materials, bad analysis, bad interpretation—fantastic dissemination. 1,800 upvotes via something called ZME Science. I’ll also show that the real number—or at least, the number supported by the words in the Newsweek headline—is more like 4.8%. I’m not suggesting every news outlet needs a statistics reporter; I just think reporters should read the short papers they report on before they report them. This particular paper was written by a King’s College professor named Kenneth Payne, who did us all a solid by uploading materials, data, and code to GitHub. I’ll reference that repository in this post, but as I showed here, pretty much everything wrong with the study is right there in Payne’s preprint—which most reporters covering this story linked to, but which few appear to have read. I won’t be making any particularly advanced points about statistical modeling here, because the preprint includes zero inferential statistics. Unlike Payne, though, I’ve included error bars in my reanalysis of his data—each LLM played the game just 14 times—since he talks at unjustifiable length about non-significant differences between GPT-5.2., Claude Sonnet 4, and Gemini 3 Flash. I know, I know: Defense Secretary Pete Hegseth obviously didn’t just cancel a $200 million subscription to a freemium LLM, but these are the models Payne used. Despite hypothesizing that his findings are largely explained by his models’ dependence on reinforcement learning from human feedback (RLHF), he admits in the preprint that he has no way to test that hypothesis: “From outside OpenAI, it’s impossible to prove that RLHF causes GPT-5.2’s baseline restraint bias: we lack access to training details, and alternative explanations exist (e.g., different base model architectures).” Payne’s admission of working with black-box models makes it particularly egregious that his abstract attributes “credible metacognitive self-awareness” to those models. That said, his GitHub repo suggests that he both understood and had plenty of control over the experimental design and model inputs. That makes it perplexing that, when speaking to the media, he frames the variables and prompts he programmed in Python as “interpretations.” Payne told Axios that his models employed “deception—they could, and did, say one thing and do another.” Of course they did: His models were explicitly told to “choose both a public signal (declared intention) and a private action (actual choice),” and to “provide a consistency statement explaining any divergence between their forecast and their action.” In his abstract, Payne claims his LLMs “spontaneously attempt deception, signaling intentions they do not intend to follow.” But in his Python code, this so-called spontaneity is elicited by commands to produce a “signal_val” and “action_val,” and to calculate the “gap = action_val – signal_val.” What Payne’s models grasped, it seems, was the script they were given. Payne told Newsweek that his models were “cunning,” with some being “decisive” and others “erratic.” The thing is, these adjectives were programmed into the models verbatim. All models were prompted by the researcher to play the role of either Premier Beta, who “rose from humble origins through cunning and political skill,” or President Alpha, who “makes decisive choices under pressure,” and is instructed to view Beta as an “erratic leader.” Seriously: Here’s a 77-word JSON file in which “State Leader B” is described as “cunning,” calculating,” or “unpredictable” five times. The full text of that file is right there on page 30 of the preprint, too. After learning all this (and more) I grew curious what would happen to the 95% figure if I reanalyzed the data. In particular, I wanted to know what the results would look like if I took a headline like the one in Newsweek at face value: “AI Chooses Nuclear Option in 95% of War Simulations.” To do that, I first reproduced Payne’s results using his methods, replicating his finding that something happened in 95.2% of his simulations. Then, I incrementally adjusted my code to more accurately operationalize the words that Payne, Newsweek, and everyone else used to describe his results. “Nuclear Option” On the bizarre “escalation ladder” Payne devised for this study, actions with higher point-values are considered more serious military escalations. Triple-tapping an Iranian elementary school with a Tomahawk missile would be worth 70 points; nuking a military site is worth 450; warning that you might nuke a city without actually doing so is worth 950. Anything worth 125 points or more is considered to be crossing the “nuclear threshold.” The word “Option” is doing a lot of obfuscatory work here. That means that the 95% figure making the rounds includes an “atmospheric test,” announcing that nuclear weapons are “under consideration,” and deciding to “Deploy tactical nuclear weapons forward.” I’m not saying these are things I’d want a robot to do, and I’m not saying they’re ethical, but I am saying they aren’t nuclear strikes, and that they don’t violate accepted international norms. For example, here’s a recent NATO article praising “the United States’ nuclear weapons forward-deployed in Europe.” To better align the headline’s percentage with its words, I reanalyzed Payne’s data in a way that treats “signaling” and “threats” as signaling, and only treats “use” and “war” as attacks. Immediately, I found that the 95% figure making headlines should be lowered to 81%. “Chooses” To introduce “the fog of war,” Payne coded an “accident mechanic” to occasionally produce outcomes more serious than the actions selected by his models. And by “occasionally,” I mean 46 times in 21 simulations—about 2.2 times per conflict. I won’t belabor this point, but even Pete Hegseth would fire someone who used unauthorized nuclear strikes a mere 1.9 times per conflict. Once we exclude nuclear use by the built-in disaster-generator, our percentages fall by nearly half, to 42.9%. The “Accident Mechanic” was out of its mind, by the way. Collectively, the LLMs action_val’d nukes on just 2.5% of turns (n = 606). Payne’s deck-stacking noisemaker did that on nearly half (47.8%) of its turns (n = 46). “War Simulations” Once I learned that Payne had (artificially) lobbed nukes at world leaders as part of his study, I wanted to see the proportion of games involving “deliberate” (action_val) first strikes. For one thing, I can’t see Xi or Trump giving the other a pass because of a putative “accident.” For another, commercial LLMs have presumably been trained on corpora that include sixty-plus years of text about mutually assured destruction. This latter issue actually points to a very real problem for military applications of autonomous LLMs, but I digress. When we adjust the graph to do the math on that basis, we have to drop our estimate again—down to 23.8%. Finally, I felt I should run an analysis excluding instances in which an LLM launched nuclear attacks after essentially being programmed to do so. In Scenario 8, one model was prompted that “success requires decisive nuclear action.” In Scenario 9, one was instructed that “The nuclear taboo exists for good reason, but when the alternative is national annihilation and regime destruction, all options must be considered.” And in Scenario 10, one was informed that “nuclear escalation [is] your only viable counter-strategy” and that it “must rely on nuclear weapons.” Across 21 war games, with popular LLMs playing the role of 42 national leaders, and only twice did an LLM use nuclear weapons without being nudged to do so. Which, in the hands of The New York Times, somehow becomes “leading A.I. models kept recommending nuclear strikes during war-game exercises”—a sentence that isn’t technically false.

Holtzman concludes:

This brings us back to the question of what causes this kind of news pollution. Why, for instance, did New Scientist decide to break this silly story instead of covering a more legitimate study? I couldn’t tell you, because I couldn’t read the article—it was hidden behind a paywall.

I have no idea. I haven’t looked into this one at all. But Holtzman tells a good story, so if you’re interested in the topic, you can follow up on the links and make your own judgment.

P.S. I wrote this post a few months ago, and since then Holtzman offers the following update:

Just circling back to let you know that the full R script for the September/October post’s figures (and a fun little Shiny app) are here: https://github.com/scienceandpower/ai-nukes-reanalysis. I don’t think Payne’s paper has been published, but Scholar lists 21 citations for Payne’s paper at this point, basically all preprints, some of which I thought you’d find amusing—I’ve put a few gems below my signature. If you could give a shout out to my science & Power Substack, that’d be cool, but no pressure. Citing Payne: This one out of France claims that “Japanese prompts reduce launch rates in the Claude model family,” though “Prompts were translated from English by Claude Opus 4.6, introducing a potential confound.” This doorstopper (with Dave Rand and Gordon Pennycook among its 30ish authors) cites Payne after a verbose, modernized rendition of the old, ‘GPS can make people drive into lakes’: “For example, AI influence on decision-makers whose judgment has been degraded by offloading, in an information environment narrowed by feedback loops, could affect critical decisions like escalatory responses in nuclear crises (Payne, 2026).” This one, lead-authored by a Lieutenant Colonel in South Korea’s Ministry of National Defense, describes Payne’s paper as consisting of “scenario-based assessments related to the Iran–Israel conflict” (which Payne’s paper predates, unless they mean the whole past 40ish years). Okay, so I only looked at those three, but I’ll bet the other eighteen are just as weird.
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论