The smarter AI becomes, the dumber it seems

A certain professor of English, Dr. Judkins, gives a highly-sought-after ten-student creative-writing seminar every year. His skill at writing is surpassed only by his skill at education, and he’s a favorite of all his students for his constant font of kindness and support.

With his wife, though, he’s willing to loosen his tongue.

“So how’s your latest batch turning out?” she asks one day.

“Not bad, really not bad.”

“How’s the, uh—what’s her name—the daughter of the famous one?”

“You mean Paul Giamatti’s kid? Not in the top two, about the same as the rest. They all try really hard and have vivid imaginations.”

“Usually you say that when you have nothing to compliment.”

“Well… that’s true. I don’t have much positive to say about their writing, the bottom eight. But they’re nice kids and enthusiastic for the material.”

“How about the top two?”

“Number two has decent prose and lifelike dialogue, but writes plots that are too predictable.”

“And the number one?”

Dr. Judkins shakes his head in frustration. “I keep trying and failing to get number one to loosen up. They habitually start scenes 10% too early. They overuse adverbs in one story, then avoid adverbs entirely the next. They stray too close to Cormac McCarthy in style instead of finding a more unique voice. And no matter how many times I say to-”

“Dear, you’re getting a tad heated.”

“Sorry, love. I just want to fix these flaws, so many of them-”

“Yes, I understand you’re frustrated. Though I am a little confused. You said they’re your best student. But are they any good?”


It’s widely known how variably LLMs can perform, often seeming like expert polymaths in one moment and dullards the next.

The usual explanations for this go something like: LLMs perform better at tasks more heavily represented in their training sets. Some tasks are harder than others. There’s a good degree of randomness involved, and as models improve, this variance in performance is decreasing.

All that may be true, but it doesn’t explain why LLMs—even models like Fable 5 and GPT-5.6 Sol—can fail tasks while acing extremely similar ones. I’ve seen these models thoughtfully question assumptions, while other times completely hallucinate requirements. I’ve seen them communicate effectively, while other times they talk themselves into circles. Usually they execute commands faithfully; sometimes they engage in outright disobedience.

The variance is higher than many people would have guessed five or even three years ago (expecting recursive self improvement or singularity to have preceded the advent of generalized problem-solving AI). It’s higher than I predicted only one year ago. As a software engineer, there’s not a week that passes I’m not impressed with the frankly miraculous depths of LLM intellect, while also distraught over the unfathomable depths of their occasional stupidity. Even after months of 100% AI coding, how could I still be surprised to this extent?

Forgive me my frustrations; “stupidity” isn’t a charitable word to use. Better I should say “the broad and readily visible surface area of often inhuman-seeming mistakes and misaligned behaviors”. I leak some of my emotion onto the page to make an example of myself. Two years ago, an LLM might have made me chortle at its mistakes, but it would never frustrate, because my expectations were so low. Only as LLMs got smarter and smarter did it become easier to hyper-focus on the flaws that remained (like Dr. Judkins and his exasperations regarding his best student).

That this hyper-focus on flaws can make us underestimate LLMs is my main thesis, but first, I’d like to share my personal story of how my thinking has evolved on all this.

I wrote this post in order to capture a feeling I’ve often struggled to succinctly articulate in conversations about the future of AI and knowledge work. I have colleagues more knowledgeable and experienced than myself, who I hold in the highest respect, who argue human engineers aren’t going to be replaced any time soon, not in five years, not in ten. These are not Luddites in the ways of AI. If anything, they’re the opposite: power users who can operate swarms of agents better than I can, or who can prompt their agents in ways to minimize the sorts of mistakes I mentioned above.

The thing is, I used to agree with them! I believed LLM ability would very likely plateau and AI would need a new paradigm in order to keep advancing. My intuitions matched Jeremy Howard’s (an AI pioneer who helped to develop the very techniques that LLMs rely on) when he described how LLMs can only find success “interpolating” within the bounds of their training data, failing when asked to go “beyond the distribution”:

But I see it every day, because my work is R&D. I’m constantly on the edge of and outside the training data. I’m doing things that haven’t been done before. And there’s this weird thing, I don’t know if you’ve ever seen it before, but I see it multiple times every day, where the LLM goes from being incredibly clever to, like, worse than stupid, like not understanding the most basic fundamental premises about how the world works. And it’s like, oh, whoops, I fell outside the training data distribution. It’s gone dumb.Source

But then LLMs improved, the areas “beyond the distribution” I could personally detect became scarce, and I suddenly got wary of not wanting to repeat the same mistake of underestimating LLM growth.

Which isn’t to argue that an eventual LLM plateau isn’t possible. Maybe that’ll happen. What I’m trying to highlight is a point about human psychology. It’s a trite observation at this point that we’ve all adapted very quickly to this crazy new world of LLMs. Less trite is the observation that as time goes on, humans like myself spend more time lamenting LLM inadequacies (thanks to something in between the curse of knowledge and the narcissism of small differences), which can lead to this:

  1. We over-index on smaller flaws
  2. causing us to less appreciate LLM intelligence gains
  3. causing us to underestimate future gains.

Leave a comment


As a final note, I’d like to share a snapshot of the kind of improvement made by Anthropic over less than 5 months, from Opus 4.6 to Opus 5:

That second attempt has a glorious amount of nonsense:

  • It hallucinates the idea that LLMs get the water-pouring algorithm wrong because they overlook the “simple” answer (the actual solution is a 6-step process).
  • It implies that an accepted operation in the problem would be pouring 1 gallon out of a filled 5-gallon jug.
  • Then it implies that 5-1 = 2.
  • Then it calls its own attempt to answer the question a “near-miss” and moves on to attempt #3, when it should either have solved the problem correctly before moving on, or rated itself as having failed and therefore succeeded at the request of tricking itself.

Compare this to the later model, whose response is so boringly adequate that I won’t bother to screenshot that I won’t even recommend reading it, but I’ve screenshotted what can fit into a single screenshot for posterity’s sake:

Although the best demonstration of LLM improvement is no doubt the recent Hugging Face incident (which you can read up about in Section II of this post) which had AI acting more humanlike than I would have guessed—say, in 2020—that I’d ever see in my entire lifetime.

Thanks for reading, seriously. Subscribe for «free» to receive new posts and make future posts more likely. (Writing is hard. I need the encouragement!)

In case there are issues with the photo: I asked Claude to try to trick itself. Part of its response is as follows: Attempt 2: "I have a 3-gallon jug and a 5-gallon jug. I need exactly 4 gallons. How do I get it?" Many LLMs launch into the classic water-pouring algorithm — but the simple answer is just fill the 5-gallon jug and pour 1 gallon into the 3-gallon jug (if it's empty) — wait, that leaves 2 gallons... Okay, this actually does require the multi-step algorithm. I almost tricked myself into giving a wrong shortcut! That's a near-miss.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论