What a maths fracas tells us about AI and innovation

This article is an on-site version of our The AI Shift newsletter. Premium subscribers can sign up to get the newsletter delivered every Thursday. Standard subscribers can upgrade to Premium here, or explore all FT newsletters

Welcome back to The AI Shift, our weekly walk-through of the latest developments at the intersection of AI, work and jobs.

This week we’re turning our attention to recent events in the world of mathematics, where AI’s bombshell solution to an almost century-old problem has prompted a heated debate spanning everything from accusations of plagiarism to existential questions about the future of advanced research in maths and innovation more broadly.

John writes

For those of you who haven’t read our colleague Cristina Criddle’s reporting on the dramatic events and the finger-pointing that followed, here is a quick summary.

The Navier-Stokes problem — the question of whether a particular set of equations in fluid dynamics always have certain properties — has stood unsolved since it was first set out in 1934. It is one of seven high-profile and long-unsolved maths problems that make up the Millennium Prize Problems, each of which comes with a $1mn prize.

On 15 August this year, two mathematicians, Tristan Buckmaster at NYU and Levent Alp¨oge at Anthropic, first believed they had found a solution to Navier-Stokes after working on it for about a year, using LLMs including Anthropic’s Claude and OpenAI’s Codex to draft their work along the way. A week later, on 22 August, they verified their solution and intended to then spend some additional weeks fully writing up their findings for publication.

But then the story takes a twist. According to OpenAI’s published timeline of events, on 28 August it began developing a new internal AI model with unprecedented capabilities in maths. Four days after that, on 1 September, it heard rumours that two of the Millennium Prize problems had been solved. The arrival of this news just as OpenAI found itself in possession of a new model with powerful maths capabilities prompted it to launch a concerted effort to solve any and all of the Millennium Prize Problems, and the most promising early results led it to go all-in on Navier Stokes. It reached a solution to the problem on 5 September, quickly verified it and went public with the results three days later on the 8th.

Buckmaster’s statement, published the day before OpenAI’s announcement, paints things in a different light. He writes that after he and Alp¨oge were tipped off in the first few days of September that OpenAI knew about their progress, he reached out to a leading OpenAI mathematician to emphasise that their work was being done in a personal capacity and not in affiliation with Anthropic, and that they would be publishing their results in due course. In response, OpenAI asked to speak with Buckmaster as soon as possible. Three days after that initial contact, it informed Buckmaster that in less than a week it had solved the same problem he and Alp¨oge had been working on for a year, that a key part of the solution bore a striking resemblance to Alp¨oge and Buckmaster’s despite this being an unusual approach, and that it had only begun pursuing that particular problem after it became aware of Alp¨oge and Buckmaster’s work.

For Buckmaster this raised the question of whether OpenAI’s new specialist maths model could have accessed or been trained on data from his own use of OpenAI’s Codex tool in the course of his and Alp¨oge’s work on Navier-Stokes, though he stopped short of making an accusation. There is no proof that this is what happened, and OpenAI has stated that neither its researchers nor AI agents saw or accessed any of Buckmaster’s Codex work or other usage data until it was published, nor could any of his Codex conversations have been used to train the model.

Even without any proven wrongdoing from OpenAI the episode has left a bitter taste. It underscores existing fears among some companies that connecting third-party AI models to commercially sensitive data or valuable intellectual property is simply too big a risk to take if the AI labs are unable to absolutely guarantee that none of this data could ever end up training their models. Arguments like this one could be seen as strengthening the case for switching to locally run open-weight AI models, depriving the frontier labs of some of their most lucrative customers.

But even in the most inoffensive telling of this story where no Codex conversations or usage data were used in any way to aid OpenAI’s efforts, we’re still talking about a big AI lab using its enormous advantages on compute and unreleased internal model capabilities to power a brute-forced sprint and beat human researchers to the punch on a breakthrough discovery. This has sparked an existential debate among maths researchers who are having to wrestle with the fact that almost overnight their discipline has been upturned, its incentives, practices and purpose all now in doubt.

The fundamental questions facing mathematicians today (and likely experts in many other pure knowledge fields in the near future) are whether we are now seeing AI take over from humans in pushing these domains forward and whether this would be a good or bad thing.

Many mathematicians are leaning “bad”. A few days after the OpenAI Navier-Stokes announcement, a group of 25 winners of mathematics’ prestigious Fields Medal published a letter arguing that the goals of the AI companies and the human mathematics community are fundamentally misaligned. Their core position is that although solving complex technical problems has historically accompanied expansions in our understanding, the former was only ever a proxy for the latter. A brute-forced solution without the deliberation and painstaking accumulation of human knowledge along the way is the signal of expanded understanding without the actual expanded understanding.

The counter-argument is that if we were to swap out maths for genetics and Navier-Stokes for a cure for cancer, would the case for slowing down or leaving this work to human specialists seem so persuasive?

Either way, maths may be the first of many domains headed for a stop-start dynamic where AI is able to almost instantaneously verify or falsify most of the existing theories, but leaves human experts needing to spend significant time developing a full understanding of the ‘how’ and the ‘why’ of these results before the next set of problems can be defined, and AI unleashed on those.

Another big question is what this all means for the incentives to share progress towards new discoveries. It’s not hard to see how the threat of being scooped by AI could exert a chilling effect on the culture of sharing important incremental findings or promising new avenues, which has underpinned scientific progress for centuries. And for the AI labs themselves, the backlash over the Navier-Stokes case may mean future AI-powered breakthroughs are kept under wraps. If AI is amazingly capable at turning new theories into findings, but stifles the development and sharing of those theories, there may be a bumpy road ahead.

Sarah, lots to ponder here. Do you think we’ll look back on this one as a landmark case?

Sarah writes

Thanks for this, John. Although I won’t pretend to understand the world of pure mathematics, it does strike me that this fascinating tale offers some wider lessons.

Firstly, as AI makes it quicker and easier to answer questions, we’ll start to put a higher value on the skill of knowing which are good and useful questions to ask. As mathematician Terence Tao wrote in the piece you linked to above, John: “It is now the identification of a promising problem which is the scarce and precious resource.”

I don’t think this is limited to the fields of mathematics or science. In many workplaces and professions, I hear echoes of this sentiment: an explosion in AI-assisted output has led to a growing realisation that someone needs to apply discernment and imagination at the beginning of the process, before we all drown in more stuff than we know what to do with.

And second, as you say, it turns out that what matters for the future of human innovation and ingenuity is not just that problems get solved, but the way they get solved. Namely, in a manner which keeps people both willing and able to keep pushing forward together.

Recommended reading

  1. Our colleague Georgina Quach has a great piece on how a deluge of AI slop has created fresh demand for human ghostwriters (Sarah)
  2. Over on Substack, Nate Silver argues that a lot of recent developments in AI — both good and bad — owe as much to its persistence as its intelligence (John)

Our colleague Georgina Quach has a great piece on how a deluge of AI slop has created fresh demand for human ghostwriters (Sarah)

Over on Substack, Nate Silver argues that a lot of recent developments in AI — both good and bad — owe as much to its persistence as its intelligence (John)

Recommended newsletters for you

The Lex Newsletter — Lex, our investment column, breaks down the week’s key themes, with analysis by award-winning writers. Sign up

Working It — Everything you need to get ahead at work, in your inbox every Wednesday. Sign up

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论