Social trust and epistemology in the AI era

In response to our changing times, a coalition of mathematicians recently launched the Association for Human Mathematics. I am a member. This Association is composed of a number of working groups. These working groups will be the engine of the Association: they will design campaigns, ideate pro-human mathematical initiatives, and articulate policy proposals in the post-AI era. We currently have a backlog of member submissions to upload. I encourage you to join and become an active member, particularly if you have good technical skills with respect to IT and digital infrastructure.

It is a sign of our times that perhaps one of the best essays about AI and math was written by an anonymous account on Twitter. Notwithstanding the inflammatory title, I encourage you to read “Mathematics is effectively dead” in its entirety now.

Perhaps the most important point is articulated at the very beginning: the input and output of mathematics is discourse. Mathematical knowledge has not been successfully integrated into the discourse unless a community of practitioners has understood the relevant ideas and can employ these ways of thought themselves. We ourselves are trained on the discourse: we receive interesting definitions, problems, and research programs from our colleagues and the generations before us. Producing high-quality mathematics involves participating in this endless recursive loop of production and consumption, or in other words of speaking and listening. This is what mathematics is.

An important aspect of this discourse is understanding where new mathematical ideas come from. Here, credit enters the scene. While it is vulgar to talk about credit, accreditation fulfills a socially important role. It is valuable from a historical perspective alone to understand where techniques originated. It is useful from a scientific standpoint to understand the mathematical methodology and technical contributionswhich led to a result. It is also important for us to know what a person has actually contributed to the discourse for evaluative reasons, when we appraise people and choose whether to hire them, offer them tenure, or fund their grants. Credit is the currency of our field. Assigning value in some way is both inevitable and necessary.

Historically, the process of assigning credit has not been too difficult; we can look at the papers that a person has authored and the public talks that they have given. And when we write original mathematical work, it is easy for us to determine how to assign credit to other mathematicians. We know which ideas originated with ourselves, which techniques we adapted from the literature, and which theorems of others we directly relied on; through citation, we direct the reader to those theorems. The literature itself provides a paper trail.

Originating ideas, of course, is not the only way we can contribute to high-quality mathematical discourse. Promoting and developing good mathematical ideas, including ideas already extant in the literature, are useful contributions to mathematical discourse. Still, credit is necessary to assign for this work too.

Social trust in the post-AI era

Enter highly capable LLMs. It is now possible to push a button and produce a new proof certificate of a wide number of mathematical statements. (In some sense this was always true: you can enumerate infinitely many true statements in ZFC by just inductively applying the axioms.) Button-pushing is an option available to basically every mathematician with a laptop, as well as laypeople, industry actors, and random people on Twitter. The proof certificates themselves can be generated in a variety of fashions: by a one-shot prompt; via extended back-and-forth dialogue which may leverage genuine mathematical expertise; through a proof-search agent harness or some other completely automated workflow; and so on.

This state of affairs has a disruptive effect on social trust. Indeed, our entire mathematical discourse is likely to deteriorate. For one, a paper no longer signals that the mathematical ideas therein were generated by the author. Even when AI use is transparently disclosed, our best available option is to attribute AI-written mathematical work to the model itself, often by naming some proprietary model in an AI use acknowledgement. This is not entirely satisfactory when the model itself was trained on the entire mathematical literature and employs ideas and high-level heuristics from this corpus. It moreover appears possible that these models may resurface mathematical ideas which were privately fed into them in personal chat logs. (For this reason, among other data privacy concerns, I believe all LLM users must switch from proprietary models to locally-hosted open source models as quickly as possible.)

The situation is worse yet. When I read a paper, even one with an AI use disclosure, I have no guarantee that the mathematical knowledge therein was effectively transferred to the author. A problem may have been “claimed”, in our ongoing ownership model of mathematics, but that does not mean any human mathematical understanding was created. In this way, dead zones in human mathematical understanding are likely to emerge unless we as a community impose other filtration mechanisms. One can now upload a PDF to ArXiv without being able to understand any of the results therein. This is not a responsible contribution to the discourse; it is something like a deepfake, a certificate of human mathematical understanding that doesn’t actually exist.

It is easier than ever to now generate a large number of proof certificates by pressing buttons. But it is not easy to reasonably integrate the results into the communal epistemological process which generates mathematical understanding. The objective of a scientific community is not “mere” knowledge; the goal is good discourse, which demands the integration of techniques and high-level heuristics into the workflows of living practitioners. In the current regime, a machine-generated proof seems to rise to the level of an empirical observation. A Lean-verified proof certificate which no human being has verified has a curious epistemic status; I think it is legitimate to argue that such theorems have not been “proven” at all. Until human beings have accepted and internalized the contents, little has been offered to the discourse, nor has value been added to the community. I suppose this is a rather constructionist perspective, but I am not a Platonist.

It is now possible to produce mathematical artifacts without participating in any mathematical activity. Button-pushing, to put it frankly, does not seem to be a mathematical activity. Running agent harnesses on the literature does not seem to be a mathematical activity. We have to radically change how we assess any mathematical artifacts as a result.

Rejecting the digestion regime

In response to this state of affairs, a scenario is frequently put on offer: an environment where people do not autonomously generate research but instead collectively “digest” and disseminate AI-generated mathematics. I believe we should reject this scenario.

It seems likely to me that we may already be near the limit of what a practitioner in a mathematical community can reasonably “digest”, and this is when most proofs are human-generated. In a scenario in which LLMs can produce technically correct proof certificates at a rate 10^5 faster than human beings, the problem of digestion becomes a serious one indeed. Even the problem of proof selection among a plethora of AI-produced proof certificates is overwhelming. Without intervention from above, human-generated mathematical contributions will inevitably be drowned out by a torrent of AI-coauthored work.

Even if we could keep up with the rate of scientific production, digestion alone does not produce knowledge. Reading is shallower than writing, and mathematical discourse cannot be predicated on a model where human beings listen and machines talk. Producing mathematics—original or otherwise—generates understanding in our community. Passive learning is ineffective for our undergraduates; there is absolutely no reason to think that bad pedagogical outcomes would not materialize for human beings doing AI-assisted research-level mathematics at the frontier.

Thus, AI tools threaten the processes which create mathematical expertise and produce knowledge in our community. Individuals need time and space to produce their own creative mathematical contributions. In other words, they need working conditions which respect their autonomy. People also need to know that their contributions will be accepted and appreciated by other human mathematicians.

If we do not thoughtfully design our working conditions, there will be a real exodus of people from research mathematics. I do not think I would have become a mathematician if it entailed verifying and cleaning up LLM-written proofs all day. That appears to me an exhausting and inefficient model of knowledge-production, a paradigm of learning math that is fundamentally reactive rather than generative. Nevertheless, this caricature has been on offer in numerous essays.

We should not be debating whether AI can do math, any more than we should ponder whether a telescope can really see Jupiter’s moons. The better question is whether AI tools can be meaningfully integrated into our knowledge-production as a community, or whether AI tools will disrupt, harm, and mitigate that knowledge-production. In answering this question, we should pay attention to what is empirically happening on the ground rather than exclusively contemplating our favourite science-fiction scenarios. Take Lean as an example. The promise of computer-generated Lean certificates was that they would empower us to fix extant errors in our literature and proofread our own work faster. In practice, we see Lean certificates employed as justification for proofs where the author cannot really explain any of the mathematical ideas. These tools are already being organically misused by various actors, to the deterioration of mathematical discourse. Within mathematics, AI tools have created a severe misalignment problem and broken our previous incentive systems. Without robust action, AI paves the way for human disempowerment.

The deterioration of open science

An interesting and well-written essay by Benjamin Antieau contrasts two modalities in the post-AI era: fast and slow mathematics. Fast mathematics might entail ambitious AI-assisted projects to understand vast mathematical landscapes; slow mathematics might involve the more traditional exploration process which we associate with building mathematical understanding. Antieau also formulates reasonable collective standards for our community, the vast majority of which I agree with.

The issue is that these two modalities cannot coexist functionally unless AI use is severely limited in scope. By limited in scope, this could mean that AI use is confined to certain problems, which are claimed in advance for certain research programs, or limited to certain modalities (like literature search and proofreading). Although I argue against AI uptake where possible, I recognize that some mathematicians will use these tools. We therefore have a social coordination problem to solve collectively.

The problem is the following. It is inefficient for mathematicians to be doing ‘fast’ and ‘slow’ math in the same domains in parallel. I do not want to be spending 6 months working on a problem if tomorrow it is one-shotted by someone using ChatGPT and dumped on the ArXiv; this kind of thing seems like a waste of my time. It is like someone has spoiled a movie, except it might also affect my career.

This is not an argument for the value of originality or the continuation of previous norms which prized being the first to get through the gate. Nor do I believe that AI-generated proof certificates “taint” or “devalue” a result; there is still value in a human mathematician proving a result even if they are not the first. Every proof method is different, and the human generation of research mathematics is tied to the production of certain goods (human mathematical understanding, novel theory, good prose, etc.) whose creation is decoupled in AI-generated mathematics.

Nevertheless, the current breakneck environment makes my working conditions significantly more unpleasant. In fact, it makes it harder to think, and to engage in deep mathematical work. Other researchers I talk to have the same complaint: this AI-accelerated environment is immensely distracting to the work of doing research mathematics. It disincentivizes mathematical thought. We may therefore require initiatives which allow us to parallelize mathematical effort and prevent the situation in which an AI agent solves a problem which a human mathematician is working on concurrently.

In the interim, this state of affairs is doing untold damage to open science. A culture of secrecy is organically emerging as people try to prevent themselves from being ‘scooped’ by AI users. Interesting conjectures are no longer stated outright in preprints. High-level proof sketches and work-in-progress are held under lock and key. A professor friend tells me he is worried after seeing the main result of a paper he is almost done writing posed as an open problem in a preprint; he feels pressured to rush out his paper now. He told me, “We have to act like there’s a script at all times which sees a new open problem stated in the literature, scrapes it, and then instantly attacks it with AI agents.”

This may sound paranoid. Actually, this concern is not so abstract when there are databases devoted to collecting and tagging all open problems in the literature, the contents of which are then offered over X to employees of AI companies after they describe their desire to “enumerate all open problems - that appeared in a paper - starting with the named conjectures” so that they can then “run it through AI agents.”

Fast mathematics wreaks havoc on the working conditions of slow mathematicians. Social coordination is necessitated as a result. Until this state of affairs is fixed, open science will close off, and social trust will deteriorate.

Demonstrating our understanding for the rest of our lives

We are not the first community to face this sort of problem. For instance, right now one can generate a good deal of images tailored to almost any specification one wants using AI tools. Nevertheless, artists have developed robust professional norms around AI-produced art, and AI-generated art is evaluated differently within the art world than human-generated work. This is true even though, for them as well as us, AI use is essentially unenforceable, and it is impossible to independently ascertain the origin of many artistic pieces. Artists have held the line, and concerted anti-AI public signaling and an honour system has shaped artistic discourse for the better.

We mathematicians are in a similar boat. It seems likely that in the post-AI era, we will be demonstrating our understanding to each other for the rest of our lives. But this is not an impossible task. We can prove our mathematical understanding through in-person interactions, conferences, seminars, and talks. We can document our mathematical process to back up our AI use disclosures (or lack thereof). We can require that recorded talks accompany every new preprint and journal submission. We can stage mathematical defenses of our work.

A recent experiment by Nihar B. Shah, an editor-in-chief of Transactions on Machine Learning Research (TMLR), might reveal one path forward: simply ask authors to explain their work. Such endeavours are undoubtedly time-consuming, but they may improve our mathematical discourse for the better. Nevertheless, we should recognize that such a policy entails an intractable workload for journal staff and the mathematical community at large unless we rate-limit scientific production. I believe this is the most urgent proposal that our community has yet to adopt.

A note of hope

Finally, let me say one thing. I am optimistic that we can redesign the incentives in our community to encourage good mathematical discourse and improve our evaluation systems. It now appears more critical than ever that we embrace a pluralistic value system for assessing new contributions to the literature, rather than valuing proof certificates for no other reason that they solve problems which people have been unable to solve before. High-quality mathematical prose, rich connections to other problems that people are actively thinking about, and deeply original ideas are worth being prized.

It appears that we, as a community, may have over-optimized on proof production before the introduction of LLMs. Now it is clear to more of us that mathematics is not only about proof production. If we continue to value problem-solving above all else, fertile mathematical terrain may well be stripmined, and mathematics becomes a race to the bottom. Rosy philosophical statements of value are not sufficient, however. We need to get serious about policy proposals that protect working conditions for AI-eschewing mathematicians and develop evaluation metrics that improve our mathematical discourse for the better.

Further reading:

Acknowledgements. I am grateful to Dylan King for editorial feedback on a draft of this essay. All opinions expressed above are solely my own and do not represent those of any employer or professional association to which I belong.

I am reminded of Thurston’s observation that mathematical knowledge can be transmitted amazingly fast within a subfield, whereas communication is slow and inefficient between subfields. To wit:

Mathematics in some sense has a common language: a language of symbols, technical definitions, computations, and logic. This language efficiently conveys some, but not all, modes of mathematical thinking. Mathematicians learn to translate certain things almost unconsciously from one mental mode to the other, so that some statements quickly become clear. Different mathematicians study papers in different ways, but when I read a mathematical paper in a field in which I’m conversant, I concentrate on the thoughts that are between the lines. I might look over several paragraphs or strings of equations and think to myself “Oh yeah, they’re putting in enough rigamarole to carry such-and-such idea.” When the idea is clear, the formal setup is usually unnecessary and redundant—I often feel that I could write it out myself more easily than figuring out what the authors actually wrote. It’s like a new toaster that comes with a 16-page manual. If you already understand toasters and if the toaster looks like previous toasters you’ve encountered, you might just plug it in and see if it works, rather than first reading all the details in the manual.People familiar with ways of doing things in a subfield recognize various patterns of statements or formulas as idioms or circumlocution for certain concepts or mental images. But to people not already familiar with what’s going on the same patterns are not very illuminating; they are often even misleading. The language is not alive except to those who use it.

I do not enjoy this process of stratification and evaluation any more than the next person. Still, some process like this will occur so long as public funding for science is low enough that there are many more people who want to be mathematicians than there are funded positions.

For instance, in my field, we tend to recognize that cost was developed into a full-fledged theory by Damien Gaboriau in [Gab00], even though credit is also due to Levitt in [Lev95], who introduced the definition of cost.

Here I mean constructionism in the philosophy-of-science sense, not constructivism in the mathematical sense. I accept proofs by contradiction.

I borrow this estimate from the excellent essay “Why do we do astrophysics?” by David Hogg.

Of course, it is also worth considering the severe ethical and environmental issues associated with these tools, but these topics are worth another essay.

Similarly, mathematicians could do a lot better by reading literature from philosophy of science, epistemology, and economics, rather than solely consuming the opinions of other mathematicians.

Above I talk about spoiling a film. Perhaps a better metaphor for intentional AI-assisted “scooping” is the following: suppose I have written a long novel for 6 months on my own. I then describe the plot to another person who uses AI tools to generate a simulacrum of my prose at length, in order to attain credit for the story before I publish it. This is exactly the kind of scenario that Antieau directly discourages in his fourth collective standard.

This kind of environment does not encourage the production of high-quality mathematics, and a similar rush-to-publish was on display in the Navier-Stokes affair, which resulted in (among many other things) some absolutely terrible mathematical prose.

Given that I spent six to eighteen months writing any given paper, it does not seem burdensome to me to talk about my work with journal editors for an hour.

In recent days, the following words of Grothendieck seems remarkably prescient:

We therefore seem to be approaching the risk of eradication, within each individual, of not only the memory of work carried out close to the source, of a “feminine” nature (often derided as “muddy”, “sluggish”, “inconsistent” - or at the other extreme as “trivial”, “child’s play”, “long-winded” …) but also the loss of this very work and its outgrowths - when this work is where novel notions and visions are conceived, grow, and come to life. Such an event would lead us to an era during which the practice of our art will be reduce to arid and vain demonstrations of cerebral “weightlifting”, and to intellectual bidding wars towards the “breaking open” of competition problems (“of proverbial difficulty”) - an era of sterile, feverish, “super-macho” hypertrophy following more than three centuries of continuous creative renewal.

This is sourced from a partial English translation of Récoltes et Semailles, which is available here. I cannot find the translator’s name anywhere. Please comment below or email me if you know the translator.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论