What can language models teach us about understanding?
Crossposted from my website. Written by me and edited in collaboration with Claude Fable (Anthropic).
This post is mainly aimed at mathematicians, as a call to think expansively about what mathematics could look like. I gesture briefly at building a new mathematical theory of 'intuition', doing for that word what the nineteenth and twentieth centuries did for 'reasoning' and 'computation' through the language of mathematical logic, lambda calculus and Turing machines. I predict that such a 'calculus of intuition' will have ramifications far, far beyond mathematics.
I have considered myself a mathematician almost as long as I can remember. Mathematics felt like a door into reality, the free communal property of humanity and yet accessible only to the few who make it their home. For thirteen years my life centered on mathematics, and it did feel like home. My PhD, finished in early 2023, went exceptionally well: I had learned an enormous amount, proved a few pretty things, established a minor reputation in my field and been offered four years of postdoctoral positions at prestigious institutions. I seemed set up for a life in mathematics and academia, and yet a few months later I quit number theory to think about cognition and AI instead. I want to explain why, what I learned about the nature of intelligence in the three years since, and most of all what my experience might suggest about new, and very different, futures for mathematics.
In 2023 I had a strong sense that within two or three years, the descendants of ChatGPT would be better at many aspects of mathematics than me and every other mathematician in the world. This had not been my view until then, and taking the prediction seriously forced me to re-evaluate many fixed beliefs. It felt important to understand how our collective understanding of cognition could have missed the language model revolution so badly. Even if my only interest were mathematics, understanding how language models conceive of it promised a deeper understanding of mathematics itself. That, coupled with the deep uncertainty about what such a transformation would do to human society, gave me a strong pull towards questions about the nature of intelligence.
Three years on, the prediction seems to have panned out. Numerous long-standing conjectures have fallen to language models in the past few months, and more fall every day. The models are far from perfect, and far from dominating all human mathematicians intellectually, but they are clearly superhuman at some aspects of mathematics. They also process mathematics differently and have different strengths today. The typical proof discovered by a model, even the resolution of a long-standing open problem like the unit-distance conjecture, feels like a very clever combination of existing ideas rather than a bold new one. To use an imperfect analogy, their superhuman performance today feels more like Deep Blue than AlphaZero.
Why are language models able to prove difficult theorems?
By the mid 2000s, chess engines beat humans using hardcoded heuristics amplified by massive computation. They won handily, though their shortsightedness left specific positions where humans could still outplay them. More importantly, their play felt 'artificial', and humans found it hard to learn conceptual principles from machine games. Contrast AlphaZero, a neural network trained by self-play to play Go and chess at superhuman levels. Its results were initially not much better than the earlier engines, but its play felt very different: humans could see the beauty in many of its moves and found it far easier to learn from AlphaZero-related engines than from the earlier era. However imprecise the feeling, there was a strong sense that AlphaZero had good taste in a way earlier engines did not. Language models are in some sense much closer to the AlphaZero paradigm than to search-based engines, and yet there have been few or no 'move 37's in mathematics — ideas far outside the current conceptual understanding of mathematicians that are nevertheless deeply generative once understood.
In July 2025 (On the Mechanical Creation of Mathematical Concepts), I speculated that language models might soon be superhuman at mathematics in precisely this way, distinguishing search, which explores a fixed conceptual space, from the more fundamental process of concept creation. In 'Computational Platonism', I present mathematics as a cycle between two poles. Syntax is the explicit trace: computations, geometric drawings, proofs. Semantics is the intuitive, illegible understanding those traces leave behind, a new perspective on the subject. Syntactic exploration comes first and its traces are integrated into semantic understanding. Exploring within that understanding then guides new syntactic production: conjectures to prove, computations to run, definitions that embody the learned intuition. These feed back into the semantics, and the cycle continues.
In this terminology, one might hypothesize that language models create semantic concepts through gradient descent during pretraining and reinforcement learning. That gives them a frame for exploring mathematics; asked to solve a problem at inference, they search within their existing concepts, backed by far greater computational resources than ours. What they currently seem bad at is creating new concepts at inference time from information in their context that was not integrated during training, such as computations they have just performed.
Even if language models never learn to coherently create new conceptualizations at inference the way a human can, training cycles are now so short that we should expect large jumps in capability every few months. If this view is correct, training on the mathematics now being generated should give them superhuman conceptualizations of it, and we should soon see proofs containing genuinely new insights into the structure of mathematics. Whether they will be able to make their implicit understanding explicit by creating new definitions and theories, and whether they will build new semantic understanding at inference time, remain open questions; the best human mathematicians still do both better.
Impact on the mathematical community
Even the current state of affairs has shaken the academic system, raising the questions of how to reorganize human mathematical activity and how to integrate language models into research. Vexingly, the answers almost certainly rest on understanding how the capabilities of language models might improve in the near future. Where does this leave mathematics and human mathematicians? I see this as a miniature of the fundamental question of 'AI alignment': how can we integrate artificial intelligence into human society so as to maximally increase the flourishing of society as a whole?
'AI alignment' usually means embodying human values within AI systems. The dual question, finding a place for humans in a world changed by AI, seems to me just as important, and the mathematical community is struggling with it right now. In September 2026, twenty-five Fields Medalists published a declaration titled 'A Severe Misalignment of AI in Mathematics', objecting to the treatment of open problems as benchmarks; their 'misalignment' is the gap between the goals of AI companies and those of the mathematical community, which is the dual question seen from inside mathematics. Mathematics is a small society but a vital part of how we conceptualize the world, so integrating artificial intelligence into our relationship with mathematics would be a strong start on integrating it into our collective relationship with the world. What might mathematics look like once we have digested the presence of an artificial intelligence that can do mathematics?
The declaration's authors, and many others, argue that the end goal of mathematics was never the production of proofs but the cultivation of mathematical understanding, shaped by the human struggle with the subject. If language models can do so much of the 'production', and what we really care about is cultivating good mathematicians, then language models look like a net loss to society; some mathematicians have accordingly called for a moratorium on using artificial intelligence to prove new theorems. I have some sympathy with this view, but it is too narrow a vision of what mathematics is.
I was drawn to mathematics by how it changed me, but equally because it offered the deepest insights about reality, whether or not I understood them. I identify with the communal human project of understanding the universe, through mathematics, science and other means. It matters to me that someone has understood something, even a topic I have no interest in. All that is essential is that this someone and I be mutually intelligible, given enough time, effort and interest on both sides. So this side of me rejoices when AI systems find beautiful proofs and make surprising discoveries.
Soon, however, there will be a flood of purely AI-generated discoveries that no human has time to digest. What would it mean for us to integrate them into our understanding? History offers a parallel. Almost two millennia ago, Diophantus posed arithmetic equations whose solution would cultivate understanding of numbers; for a long time before and after this, such computational problems were the touchstones against which mathematicians sharpened their understanding. As facility with computation grew, and especially once computation was automated in the twentieth century, mathematicians gave proofs a more central place, to explain the volume of computations they were generating.
A proof can be seen as organizing large swaths of computation at once: by distilling the unifying essence of a family of computations, it structures the mathematical landscape. The axiomatic method, first in geometry and then across mathematics, can in turn be seen as a study of the syntactic structure of proof itself, and a direct line runs from Hilbert's search for a satisfactory account of that method, through Gödel's dashing of his hopes, to Turing's formalization of computation. Proofs and computations are inherently linked, as the last century made clear; computational complexity has turned up many surprising connections between computation, provability, randomness and more.
New mathematics as a consequence of language models
Integrating a flood of new results has historically meant building a theory one level up. Proofs organized the flood of computations, and the centrality of proof in turn produced a theory of formalization, axiomatics and computation with ramifications far outside its original domain. I think we are due for another seminal step of this kind. Finding a proof is, or shortly will be, largely automated. This does not mean that all theorems will be proven, just as not all computations have been trivialized, but we can now step back and survey the landscape of proof as nineteenth-century mathematicians surveyed the landscape of computation. We need something to organize the flood of proofs and the intuitions behind them.
The Fields Medalists' declaration argues that 'solving problems is only a tool and proxy for achieving the primary goal of conceptual understanding and insight.' That conceptual understanding is what I have been calling semantics. I agree that it is crucial to mathematics, and precisely for that reason we should make it a first-class citizen of the mathematical universe.
Can we say anything useful about how semantics, the illegible locus of intuition, arises from syntax? At its most prosaic, a language model that can resolve Millennium Prize problems is a very complicated mathematical function that implicitly encodes the essence of the mathematics it can prove. One way to reach that implicit knowledge is to sample proofs, by running the model on a wide variety of questions, and try to reverse-engineer its understanding from them. Every way of transmitting human understanding has been a close variant of this, passing necessarily through syntactic traces. Might things be different for language models?
An ambitious goal is to extract semantic understanding directly from a trained model. This is what interpretability tries to do for language models in general, with no particular focus on mathematics, and the field is in its infancy. We understand very little about how such understanding should be extracted or in what form it should be represented; people speak loosely of 'representations' with no precise definition of what they are or how they relate to the 'understanding' a model encodes. Moreover, whatever is extracted must currently be expressed in ordinary syntax, natural language or mathematics; we lack a language for describing semantic understanding itself.
Syntax and semantics
That language would be something like a calculus of 'intuition'. The theory of computation formalizes specific kinds of legible reasoning well, but intuition seems to have features that resist the language of Turing machines; I explore the difference in 'Reasoning Was Not Made for Deduction', and here only sketch what such a theory might look like. Think of a syntactic space whose points are mathematical statements, in a fixed language, and whose paths are proofs or computations leading from one statement to another. We might then think of intuition as a far more local notion: it tells us which next steps look promising without necessarily having the end in sight. It resembles a 'vector field' on this space, and a language model trained in this format gives rise to such a field quite literally, a probabilistic one whose output probabilities point along the possible paths out of each statement. The geometry of this field implicitly defines a semantic space. One natural way to make that precise is to define the distance between statements A and B as where the probability is that of the model, given B and whatever else is in its context window, generating a proof of A from B. This measures how easily the model gets from B to A, not whether B implies A. The distance is not symmetric, and the degree of asymmetry itself might encode important semantic information: a theorem is 'deep' when it is near its consequences and they are far from it. Each proof path carries its own log-probability, and different models, that is, different intuitions, assign different distances. Descartes' unification of algebra and geometry is an early human example of one intuition bringing statements close that another had kept far apart. A new definition goes further: it is a new point in the space, and a good one collapses distances everywhere at once. Whether a model has created a concept, rather than merely searched, should show up as exactly this kind of surgery on its geometry.
Intelligence may have more to do with constructing and exploring such semantic spaces than with proving any particular results, and understanding their shape seems essential to extracting insight from automated provers. We should be able to compare an intuition trained on Euclidean geometry with one trained on Diophantine arithmetic and extract what they share, mirroring Descartes. We should be able to merge, combine and fine-tune intuitions through their geometry. We should be able to read off from the geometry what kinds of proofs a model is likely to find, and which not, before sampling it, and conversely to predict the shape of an intuition from the syntax it is trained on.
Two proofs of the same statement through different fields form a loop, and the loops that persist at large scale are the great unifications. What is the homology of mathematics, and how many dimensions does it have? This geometry should also naturally have a dimension of evolution through time: as an intuition is learned, distances should collapse in bursts, one concept at a time. Most of all, we should be able to turn intuition into explicit definitions, conjectures and heuristics that make it easier to express.
The value of mathematics
Where does this leave mathematics? The Institute for Advanced Study in Princeton is perhaps the innermost sanctum of modern mathematics, its permanent faculty an eminent few whose tastes often decide what is fashionable. Yet when I visited in 2024-25, Akshay Venkatesh and Govind Menon ran a seminar on older mathematicians, back to the Ancient Greeks, with the explicit goal of learning how previous generations had reacted to decisive changes in how mathematics was practiced and perceived. We wanted to know what the past could teach us about the present moment.
I came away with a far more expansive view of what mathematics had been and could be; against that history, the modern focus on proofs and axiomatics looks narrow. The axiomatic turn is an important innovation with much to recommend it; learning Grothendieck's algebraic geometry was, for me, a deeply spiritual experience. But we should not lose sight of the broader vision, or take the current state of mathematics for its canonical form. Mathematics, to my reading, has always been about finding the right language for fundamental notions such as space, time and the nature of the mind. Until very recently it was an integral part of the scientific world view, taking as much inspiration as it gave. Discoveries about the motion of the planets motivated the invention of calculus. Gauss tried to measure the curvature of space to check whether we live in a non-Euclidean world. A long history of wondering about the mind and the nature of logical operations culminated in computers and the theory of computation.
In the long view, the narrow, specialized mathematics of the post-war period is far from the norm, and the present moment may be forcing us to re-examine why we do mathematics at all. My own deepest reason is to understand the deepest questions about reality. We now have a chance to understand the nature of mind in far more precise detail, and mostly by means of mathematics: neural networks are deeply downstream of the mathematical tradition. The 'calculus of intuition' proposed here is one attempt at an effective theory of why deep learning works. Even if it is on the right track it will be one piece of the puzzle, and such investigations will probably reshape our understanding in ways no one expects, as every past moment of scientific progress has. That understanding is almost certainly a necessary step towards finding the right place for human mathematicians in the society we are building. It would be a pity if mathematicians left the most interesting scientific question of our time to others.