Mathematicians, OpenAI, and a $1m Problem

The Navier-Stokes equations mathematically describe how fluids move, and there’s been a $1m bounty on proving their solutions since 2000. Twenty-six years later, those tough math problems became the subject of a complicated discussion on ownership in the age of AI.

On September 8, Tristan Buckmaster, Levent Alpöge, and Matei Coiculescu published proofs showing that several closely related equations, including 3D incompressible Euler, can blow up in finite time when you push them with a smooth external force. “Blowup” means some quantity in the solution runs off to infinity after a finite amount of time rather than staying finite forever, which is exactly the behavior that the $1m Clay problem asks about.

They verified the proofs in Lean (a programmatic proof assistant), so the argument is machine-checked rather than resting on a referee’s reading. Terence Tao wrote up the results on his own blog, which is a fair proxy for how seriously the field is taking them.

The work took about a year, and most of it was slow going. According to Buckmaster’s account, the breakthrough came on August 15 and was verified by Lean on August 22. They used several models throughout and paid out of pocket: Anthropic’s Claude, OpenAI’s Codex with GPT-5.6 Sol, and more recently Astra for writeups and auditing arguments. Every draft of the project went into their Codex sessions.

Buckmaster says he emailed OpenAI privately on September 3 after hearing that his work had reached the company, hoping to head off a collision. On September 6 he spoke twice with Sébastien Bubeck, who leads OpenAI’s math team. Bubeck told him that an internal model had produced a roughly 100-page proof of finite-time blowup for forced Navier-Stokes, using smooth forcing under options c and d in Fefferman’s formulation of the problem.

That is the same narrow route Buckmaster and Alpöge had chosen, and it is not the route you land on in a few days by handing a model the problem statement. He was initially told that the model had received very little human input, and by the end of that same call, the narrative was different: an entire team had been working on it, the model had been warmed up on easier equations first, and the prompt he was shown had itself been written by prompting Codex.

He asked when the first prompt went out. The answer, eventually, was the past few days, after information about his work reached OpenAI. He asked whether the internal model had been trained on, or had access to, the Codex sessions holding every draft of the project. He was told the model does not look up user data. When he asked again about training specifically, he got nothing back.

Bubeck has called the allegations circulating about him false and inflammatory, says he approached the discussions according to academic norms, and has promised a fuller response. Buckmaster is careful in his own statement: he has not seen OpenAI’s proof, he does not know whether their data was used, and he says he is not accusing anyone of anything. Those qualifications are doing real work, and a lot of the commentary has thrown them away.

The authorship fight will get sorted out in public over the next week, and whether Bubeck was rude on a phone call is the least durable part of this. Two people spent a year on an obscure program, and when they needed to know what had happened to their own unpublished drafts, they had to ask a vendor and hope for a straight answer.

They did not get one, and there was nowhere else to look: no log they could read, and no way to reconstruct what had left their machines or where it went. Buckmaster has no way to find out whether a model was trained on his Codex sessions, which is why his statement stops at describing the question rather than answering it, and that same gap sits underneath a lot of ordinary software work.

The Takeaway: What you hand over without noticing

Nobody signs a document forfeiting ownership of their work. It accumulates. You put an architecture decision in a chat because that is where you happened to be thinking, an AI model reads the whole repo to answer a question about one function, and a month of design iteration on an unreleased feature ends up as conversation history inside a product owned by a company that might compete with you next quarter.

For most teams this never becomes a problem, and one research dispute is thin evidence for a company-wide policy. The exposure is uneven, though. If you are building against a well-capitalized lab in a narrow market, or if your advantage is a specific approach rather than an existing business, your unfinished work is the asset. Researchers publishing in a field where three groups worldwide are chasing the same result live in that position permanently, and so do plenty of startups.

Hedging is boring and mostly mechanical

Keeping one provider from becoming the only party who knows what you are doing takes a few unglamorous decisions.

Start with tooling you can read. When a tool’s source is public, like Kilo’s, the question of what a request actually includes gets answered by reading the code, rather than by asking a vendor to characterize it and waiting to see how specific the answer is.

Then spread the work around. Routing across more than one family of models means that no single provider holds all of it by default, and you can send exploratory reasoning somewhere different from a routine refactor. Per-token pricing at the provider’s own rate keeps that easy, because you have no committed spend making a move awkward and no plan tier deciding which model you are allowed to think with.

Using several models is not the same as spreading the work, though, and Buckmaster and Alpöge are the clearest illustration of the difference. They queried at least three, yet the project accumulated in one product’s sessions, which left a single company holding the drafts and answering the questions about them. What protects you is where the working state lands, so pick tooling that keeps that under your control rather than in whichever vendor’s chat you happened to open. None of this stops a provider from behaving badly, but it does keep any one of them from holding the complete picture.

The math here will outlast the argument about it. Before Bubeck’s fuller response lands, go look at where your own unpublished work currently lives, and count how many companies can see all of it at once; then consider choosing a tool that hedges against a closed model ecosystem.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论