Taste Is Not Enough. Reality or Bust


A paper published in Nature on March 25, 2026 describes "The AI Scientist," a system built by Sakana AI that automates the full cycle of scientific research: idea generation, experiments, analysis, writeup, even peer review submission.
The marginal cost of producing paper-shaped research output is collapsing.
So, when paper production becomes cheap, what is next?
100,000 axiom systems and counting
Stephen Wolfram spent years systematically enumerating all possible axiom systems. Each axiom system defines a "possible universe of mathematics": a different set of starting rules, a different universe of theorems. Most universes are empty or trivial. Of the ones that are not trivial, many are bizarre. Our entire familiar mathematics occupies a tiny corner of this space. Logic, specifically Boolean algebra, turns out to be perhaps the hundred thousandth axiom system you would encounter if you enumerated them by complexity.
Wolfram found nothing obviously special about the axiom systems we actually use. His suspicion (not a proof) is that we study them for largely historical reasons: they are generalizations of arithmetic and geometry from ancient Babylon. The space of possible mathematics is vast, our explored corner is small, and we are in this particular corner because of history, not because of any intrinsic property of the systems themselves.
Why do we use them? Because humans decided they were interesting.
The AI Scientist has the same problem, one level up
The AI Scientist can generate research ideas, execute experiments, and write papers. Sakana reports a cost of roughly $15 per paper. One of three papers passed peer review at a ICLR workshop (not the main conference track), with humans filtering the most promising outputs before submission. The system can produce formally structured research outputs. It cannot yet tell which ones matter.
The problem is not just quality filtering. A separate study in Nature earlier this year analyzed 41.3 million research papers and found that scientists using AI tools publish three times more papers and get five times more citations. Great for individuals. But collectively, AI-driven research covers less topical territory. It clusters around already-popular problems.
In the Wolfram analogy: a machine that evaluates "interesting" by pattern-matching against known mathematics will keep steering you back to the hundred thousandth axiom system and its neighbors. Lots of exploit, much less explore.
So what is the actual scarce resource?
This matches my own experience using AI agents for research and teaching. The agents are shockingly good at execution. Give them a clear task with well-defined scope and they deliver something genuinely useful, fast.
But "work on the next most important task" only works if someone figured out what the important tasks are. The agent does not decide which questions matter. The moment you ask it to define its own scope, you get the AI Scientist problem: lots of output, most of it predictable, much of it wrong in ways that require domain expertise to even detect.
The scarce resource is judgment. The ability to look at a vast space of possibilities and say: this one.
That is the comforting answer, anyway. AI does the grunt work. We provide the taste, the direction, the vision. We stay at the center of the universe.
Except: that story is cope. Rich Sutton's " Bitter Lesson " showed that every time researchers tried to hand-code human knowledge into AI systems (chess heuristics, vision algorithms, Go strategies), brute-force scaling eventually crushed the hand-coded approach. Human judgment about what matters may just be the next ontology in line to be bypassed. But even if it is not, history suggests it was never as reliable as we like to think.
But the world has a vote
Wolfram's enumeration is purely abstract. The axiom systems just sit there, inert. But science interacts with data from the world we observe. You hypothesize, you collect data, and reality tells you whether you are wrong. And that feedback loop has a history of promoting "useless" systems to central importance, often over the explicit objections of the people who understood them best.
Godfrey Hardy, a godfather of number theory, wrote in 1940 that number theory had a kind of supreme uselessness, that no one had discovered any warlike or practical purpose for it, and it seemed unlikely anyone ever would. And he was proud of that uselessness, as a sign of the supreme taste of a pure mathematician.
Thirty-one years after his death, RSA encryption arrived, and modern cryptography now depends heavily on the number theory Hardy was so proud to call pointless.
Maxwell predicted electromagnetic waves in 1865 as a mathematical consequence of his equations. Hertz demonstrated them physically in 1887, and when his students asked what the discovery was good for, he replied: "It is of no use whatsoever. This is just an experiment that proves Maestro Maxwell was right."
Marconi built the wireless telegraph less than a decade later.
Notice what Hardy and Hertz have in common. They were not amateurs. They understood their own discoveries better than anyone alive. Their taste was extraordinary: out of the vast space of possible mathematics and physics, they picked systems that turned out to be profoundly important. But their forecasts of usefulness were completely wrong. Hardy looked at number theory and said: this is beautiful and this is deep. He was right about that. He was wrong about what the world would do with it. Hertz looked at electromagnetic waves and saw a confirmation of Maxwell. He was right about that too. He could not see the wireless telegraph.
The distinction matters. Taste selected the right systems to study. But taste could not predict what those systems would be for. That was decided later, by technologies and applications that did not yet exist. The world retroactively decided which "useless" formal systems had been important all along.
So the comforting story ("AI does execution, we provide the visionary taste, we stay at the center of the universe") is incomplete. Taste is real but taste without reality is flying blind. Entire fields operate this way: elegant theory frameworks that survive for decades because they never invite reality to correct them. And the people with the best taste in history still could not see where their work would land.
The right question is not "who has the best taste?" It is "what kind of feedback loop lets reality surface the value that taste alone cannot see?"
Can AI close that loop?
In some fields, it already has. 20 years ago, Mechanical Turk returned human judgements through API calls. Now, autonomous wet labs ( Emerald Cloud Lab, Strateos, RAPID-200 ) accept experimental protocols via API and return physical results without human hands touching anything. An AI agent can already design an experiment, submit it to a cloud lab, and get data back. The loop with physical reality is not a future idea. It is existing infrastructure.
And still, the narrowing problem persists. Ten thousand automated experiments over a weekend still require someone (or something) to decide what experiments to run. The labs automate verification, not direction. Reality is the slowest, most expensive API there is. A clinical trial takes years. Growing a test crop takes a season. AI generates hypotheses at near-zero marginal cost, but verifying them against the physical world still costs capital and time.
So, the question is what happens when AI-generated ideas start getting corrected by the world. That is the difference between an AI that enumerates the space of possible mathematics and one that discovers non-Euclidean geometry because spacetime forced its hand.