“Deep theorems were scarce and difficult and so became an effective mechanism to identify deep thought. AI has broken this system.”
[This is a guest post by Bryna Kra. This blog post was initially written in a different file format and converted using AI. — T.]
For generations, mathematicians have treated the production of theorems as a clear measure of success. The stronger the theorem, the deeper the proof, the more surprising the connections, the greater the achievement. Positions and prizes are based on these theorems and the mathematicians making the breakthroughs set the directions for future research.
But theorem production was only part of what we cared about. It was a proxy for something harder to measure: understanding. A major breakthrough meant that years were invested in learning a subject and uncovering hidden aspects, accompanied by work to make the answer apparent to the community. Deep theorems were scarce and difficult and so became an effective mechanism to identify deep thought.
AI has broken this system.
Artificial intelligence has lowered the cost of producing sophisticated proofs. As the models improve, the list of deep conjectures that fall will grow. Someone with little knowledge in a field can now generate a manuscript that reads like a polished article, cites the literature, and combines techniques that would have taken years to master. The machines are producing solutions faster than the mathematical community can read them, much less digest them.
This week the changes arrived in my field. The Nivat conjecture is a beautiful problem at the intersection of combinatorics and dynamical systems. Imagine an infinite grid of square tiles, with each tile colored either red or blue. Look at every rectangular window of a fixed size, say tiles wide and tiles high, and count how many different color patterns show up in that window. The conjecture says that if there are at most such patterns, then the entire coloring repeats under some fixed, nonzero shift of the grid. In other words, some precise limit on local behavior gives rise to simple global behavior, periodicity.
Though this statement can be explained without formulas or higher mathematics, the conjecture resisted our efforts for nearly three decades.
In work that we first circulated in 2012, Van Cyr and I proved a partial result, showing that periodicity holds when the number of patterns is at most half the area of the window. We translated the combinatorial question into one about the geometry of dynamical systems, studying ways in which information in the system does (or does not) propagate. Progress on challenging problems does not happen in a vacuum, and our work began by studying and building on earlier results in the literature [1, 2, 3]. Soon after, Jarkko Kari and Michal Szabados developed a different approach, viewing the combinatorial question as an algebraic object, using certain types of polynomials to find the infinite patterns.
To those of us who worked in the area, this was a tantalizing situation. Two different methods to approach the same problem seemed to shout to us: find the conceptual bridge between them. We tried, without success, and moved on to other problems. Colleagues continued pushing the boundaries of what we knew, combining the varied techniques in increasingly powerful ways [1, 2, 3].
But synthesizing different approaches is one of the ways in which frontier models excel. A machine does not need years to absorb distinct areas of mathematics before it can find the connections. It can translate notation, compare partial results, and try thousands of combinations without being discouraged. The Nivat conjecture has been proposed for inclusion in Google DeepMind’s Formal Conjectures project, turning the problem into a target for automated reasoning and making more people aware of the question.
The manuscripts have started to arrive. This week I received several purported proofs of the full conjecture from researchers making their first foray into the subject. Some acknowledged using AI, but only for polishing the English or checking the proof. When I asked if the authors would meet over Zoom to explain their arguments, none accepted the invite.
Silence does not mean that the authors used AI and a proposed proof is not a theorem. But merely disclosing AI use and checking correctness of the proof, both of which are essential, miss the point. Suppose these results are correct. That is a mathematical contribution. Yet it opens more questions. What did we learn? What did the authors contribute? What impact will this result have?
A proof is more than a certificate that something is true. Instead, it is a story, a picture, an insight, an explanation. A proof highlights novel ideas and opens new directions for what we should ask next. It becomes part of the toolkit of the community. A deep theorem changes how we think, not because of its statement, but because of what it teaches us. As Bill Thurston wrote on MathOverflow in a 2010 response to a question about what mathematicians do: “The product of mathematics is clarity and understanding. Not theorems, by themselves.”
These new texts certainly contain knowledge, just as a data file of bits of 0’s and 1’s contains information. But human authorship is more than the string of words that prove something. It is intellectual responsibility, understanding of a new concept, the ability to respond to questions, the sorting of new ideas from existing scaffolding. And perhaps most importantly, it is incorporating the result into the corpus of our understanding.
I welcome AI-generated proofs and believe that human-only proofs will become rare. This is a profound change of perspective, but mathematicians have always used external tools and machines. AI is already a powerful mathematical instrument and we, as a community, need to incorporate it responsibly into our practice of mathematics. The question is not if machines should participate in discovery. They already do. The question our community needs to address is what we value in our mathematics now that certain kinds of discovery are no longer a scarce resource.
The tradition of journals publishing novel results, hiring committees rewarding those publications, and prize committees recognizing the person who completed the result do not suffice. Peer review is a slow process that relies on the unpaid expertise of a small group of overworked researchers. Any single submission can be handled, but the current volume is drowning the reviewers. The careers of young researchers depend on producing theorems, with incentive to produce as much as possible as quickly as possible. The time for deep reflection, necessary for deep understanding, does not happen when someone types the question into a model and quickly produces a manuscript.
All of these issues existed before AI entered our field. Unfortunately, we no longer have the luxury of addressing them in a leisurely manner.
Mathematics is at a watershed moment. Our incentive structures are misaligned with the prolific output of highly accessible AI models. Careful verification and stellar exposition will not happen when there is no reward for those tasks. Producing a paper is no longer enough: authors must be able to explain the proof’s mechanism and how they arrived at this point. Journals need to distinguish and credit the roles of discovery, proof, formalization and explanation. Exposition must become an important component of any major intellectual achievement. We advance mathematics not by what we write, but by what we learn, check, apply, and teach others.
Finding the right concept, crafting the right definition, and formulating new questions have always been part of our intellectual work. The creation of collective understanding, the ability to communicate ideas in a lecture, and the sharing of ideas informally that spark research have always been admired. Such achievements are harder to quantify than theorem production, and so have been treated as by-products. But these are the key aspects shaping the future of our field and their value is at least as important as knocking down the next conjecture. Our incentive system needs to be realigned to reflect what we truly value, and not just what we can easily measure.
Mathematics is the canary in a coal mine for all parts of intellectual life. Its claims can be carefully checked and so we are witnessing the disruption in real time. But when AI disrupts the tried and tested standards in a field like mathematics, what will it do to fields where evidence and interpretation are more contested? Science, law, public policy, and art will all have to find ways to assess what is valuable amidst the abundant output.
The arrival of machine generated mathematics does not replace the role of human expertise. Instead it reveals why we valued that expertise. The goal of our discipline is not simply to resolve conjectures, but to enlarge our capacity to reason and open our minds to new ways of understanding. It will take our entire community to create new standards, discover new ways of creating mathematics, work with AI models to open new horizons of knowledge, and develop the tools to approach them. It’s truly an exciting time to be a mathematician.