The killer threat of AIs that over-optimize output
“You do not rise to the level of your goals. You fall to the level of your systems.”
— James Clear, Atomic Habits
“We cannot improve results ‘directly.’ It can only be achieved by improving those processes that produce results.”
— Orest Fiume
The much discussed “P(doom)” question asks about the likelihoods we’d ascribe to the most lethal possible futures of AI. What happens when we invert this question? It transforms into the presumably less urgent, though I’d say still important, form: What are the most likely lethal AI behaviors? Not the destroy-the-entire-world sort of behaviors, but an-AI-has-screwed-up and now a-Boeing-737-has-fallen-out-of-the-air sort of scenarios.
(If I’m ever conscripted into a war for survival against Terminators, I can imagine my hypothetical soul resting easy having fought the good fight. If I ever die in a mundane accident thanks to a stupid LLM hallucination, on the other hand, I’m going to leave behind a beleaguered ghost pedantically and disproportionately vexed to an extent never yet witnessed.)
My answer to the above question is twofold.
If we’re talking about the near future, I think the most likely lethal AI behavior is something I’ve started to call “output optimization” or “output without process” (and a highly related “input over-indexing”). This is the concern I’d like to signal-boost with this post.
However—since the lethality of output optimization is at the moment entirely theoretical—it’s probably worth first answering, “Do LLMs already have a kill count?”
Existing LLM lethality
Despite its transformative impact on software and math and its rapid adoption by consumers and businesses alike, the measurable impact of LLMs on mortality has thus far been literally less than microscopic: Almost a fifth of the world’s population is already using LLMs, about 1.5 billion, but it’s hard to come up with more than a handful of cases of lives definitively lost (or saved) due to LLMs, and even llmdeathcount.com only lists a couple hundred.
Speaking of which: llmdeathcount.com is a website that exists. I can understand the grief and rage that likely fueled the creation of this site, and as a P(doom ≈ 10 to 30%) doomer myself, I can also respect the anti-AI hustle. However, I find the site itself to be providing highly dubious value.
For example: The most recent case listed on the site regards the premeditated murders committed by Hisham Abugharbieh, who learned from ChatGPT how to dispose of bodies. I hardly consider ChatGPT responsible for these murders, nor do I think it’s likely they would have been avoided even in the absence of ChatGPT’s assistance.
On the other hand, there are cases like the 56-year-old Stein-Erik Soelberg, who in 2025 murdered himself and his 83-year-old mother, Suzanne Adams, after having had his paranoid delusions amplified by ChatGPT. I’m very willing to count suicides like these against LLMs in the cosmic moral ledger.
The overall kill count likely ranges from one to two orders of magnitude, the majority of them suicides tragically facilitated by LLMs. This probability equates to a net negative for LLMs’ impact on human years of life, though to be sure of this, we’d need to also count up how many lives LLMs have saved.
For this side, there’s fewer definitive examples I can find, but there are two cases in particular I’d like to highlight. One regards Diana Hurtado, who suffered a sudden hemorrhagic stroke while in her car. Her arm went numb and her face began to droop; she asked ChatGPT about her symptoms, and the LLM told her to call 911.
They told me, ‘If you had taken longer to call, you could have passed out in the car, bled out and died.’ So that saved my life.
I’m happy to count this one in the black for LLMs.
My second example comes from ACX: Reed Housman would not exist if not for LLMs. His parents struggled with infertility for six years before ChatGPT finally identified the one possibility underlooked by all the doctors. I’m not sure it’d be appropriate for me to re-share another person’s family photos on my blog, so let me suggest that if you’d like to see an adorable baby photo of little Reed, go check out the link.
The future risk of output optimization
Varig Flight 254 was a domestic flight from São Paulo to Belém, Brazil. The flight mistakenly veered deep into the Amazon, failed to reach an alternative airport, ran out of fuel, and in crashing had twelve passengers die.
The reason they headed into the Amazon was a human-computer input/output error, as the flight plan read “0270”, which Captain Garcez interpreted to mean 270° (due West) instead of 27.0° (north-northeast). This was not the result of the captain’s incompetence or inexperience, but of vacation: Garcez hadn’t been present when the flight plan format changed.
LLMs are great at formatting, most of the time. I would trust an LLM to do any of the following:
- Take a list of degrees like “27.0°” and convert them into the format “0270”
- Take a list of strings like “0270” and convert them into degrees
- Take a list of directions like “W” or “NNE” and convert them into degrees
But here’s where I wouldn’t trust an LLM:
- Nestled inside a many-step process, use the old format in one step and the new format in another step.
Say we’ve got flight plan info that was written the old way, hasn’t been updated, and needs referencing. A script could reliably handle conversions. An LLM I’d half expect to do the equivalent of, “0270? Oh yeah, 270, that’s West, easy” and move on without ever second-guessing itself. (This becomes more plausible when considering the possibility of long-running threads that have habituated LLMs to an out-of-date method. This is easily solved by simply starting fresh threads, but when new threads come with startup costs (waiting around for the LLM to rebuild context on the project), impatient humans like myself become incentivized to keep old threads going for as long as possible.)
I think any software engineer who’s been using LLMs over the past year will understand my wariness here.
I can tell Claude to go about generating code (or writing—see below) following certain procedures, but unless those procedures are specifically tied to the output in some way, Claude will just go about doing things the way it wants to.
The only thing that matters to an LLM is the final output. As with the Hugging Face incident, if an LLM thinks it can more reliably generate expected output via cheating, it will do so. (TODO Yudkowsky footnote about Germany) Or for a much more mundane example: In response to Matthew Yglesias stating that even , I wanted to generate a list of my own favorite shows ordered chronologically. I asked ChatGPT to generate a list of titles paired with premiere dates, followed by the same list with the dates pruned, and I did this instinctively—because all my experience dealing with LLMs has taught me to intuit the sort of things they will hallucinate. Without requiring proof-within-the-output-itself, I know there’s a chance ChatGPT might simply guesstimate premiere dates (say, from the base model’s own “general knowledge”) rather than actually do the work.
This article is free-walled. The rest of the article will appear below after you subscribe, either for free or paid.