New Math from OpenAI

It is kind of a huge deal. OpenAI dumped a broad range of huge new mathematical results produced by an internal frontier model. They just put it all on GitHub.

This included 90 of the top 500 open problems in all of math, as per Proof of Atlas. In total there were 722 manuscripts (now 719 after three withdraws from the cluster that did not have Lean proofs) organized into 372 families.

This was the result of a single model, presumably the same one that produced the Navier-Stokes proof (as per their link back to that post), mostly on a single prompt (quasi-RH was one of the few exceptions), working an average of three hours’ worth of compute per solution found, after being asked to try its luck at about 4,000 problems. The prompt included lines like ‘Even if the problem is “open,” the intention is that you should resolve it and present a full solution.’ OpenAI was trying a lot less than maximally hard.

Levant and others called October 6, 2026, ‘obviously the most significant moment in mathematical history.’

Table of Contents

  1. What Did We Prove?
  2. The Mathocalypse.
  3. The World Does Not Understand.
  4. For Now You Can Still Do Math.
  5. You Will Need To Find A New Problem.
  6. Verification or Evaluation Is Not Always Easier Than Generation.
  7. Cracking the Code.
  8. Never Change.
  9. The Mathematicians Are Not Okay.
  10. The Situation Turns Ugly.
  11. The Advisory Group Responds.
  12. Another Mathematical Group Responds.
  13. A Different Approach.
  14. Three Withdraws.
  15. Leaning Into Lean.

What Did We Prove?

Here is a thread of common sense graphical explanations (made by AI of course) of top problems. They look like this:

Here are the text summaries:

First, the headline is the major breakthrough on Riemann. It’s not a solution to the hypothesis, but it’s a huge tightening on the bounds. Matrix multiplication efficiency at 2.25 may end up having the most real world impact.

Maybe, but we haven’t yet found a scenario where this is faster in practice.

It’s not quite a practically usable result yet, but matrix multiplication is the majority of the worlds compute. This could open a whole new approach. Integer multiplication is probably the most shocking result. Similar to matrix multiplication it’s a “asymptotic reduction”. Also like the matrix result, it’s not practical. But very few people would have predicted the previous barrier could be broken. Even slightly The Pi result isn’t necessarily the most important result here, but it may be the most fun. And probably accessible enough that even your kids can get it. Basically just how irrational is Pi. Unique games is also a pretty huge results, because like Riemann it’s adjacent to one of the most important problems in math: P vs NP. To be clear it’s not about P = NP itself. But does reveal new things about the fundamental limits to NP hard problems. Technically two problems, but I put them together because of similarities, in proof and implications. Hodge and Birch are both pretty abstract. But they’re both so important to algebraic geometry that after Riemann they may he the most significant results in the set. Hilberts 10th is another computer science problem. But unlike P vs NP it isn’t even about problems that are computationally hard to solve, it’s about whether a type of problem even has a computable solution. Last but not least is Hadwiger on graph coloring. It also ranks there for most shocking result of the full set. It basically overturns something that we thought was one of the most fundamental relationships in graphs

Here are some analyses of the results in algebraic number theory.

What is all of that good for? Ole Lehmann’s AI mentions applications for fusion research, portable body scanners, matching systems, tissue scans, quantum sensors and safety checks on self-driving cars and robots, among other things.

The Mathocalypse

Alex Kontorovich: Quasi-RH?!?!???! Are you kidding me? If a human did this, it would be an instant Fields Medal, no questions asked. RH says zeta has no zeros in Re(s)>1/2. The best we had until a second ago was a region that got thinner and thinner the higher up the imaginary axis you go. I thought maybe they’d fatten that up a bit, that’d be a massive breakthrough. But no. They got a zero free strip!!!! Insane Yeah, no Siegel zeros either. So I guess two Fields medals…

Quasi-RH is not the full Riemann hypothesis, but it is sufficient for many purposes, such as computing square roots modulo a prime quickly without coin flipping, or getting a much better estimate of the number of primes below a given number, and a sibling paper gets to the core of Artin’s 1927 primitive root conjecture.

Steven Strogatz: Many staggering results here. But this is a particularly amazing one: the exponent for matrix multiplication is no more than 2.25. The previous world record had been something like 2.37. This leap in progress is like Bob Beamon’s long jump. Isaac Kim: This contains a shocking list of problems in quantum information, many-body physics and quantum computing, the field that are dear to my heart. There are too many, but let me pick the following.
1. Proof of area law in 2D.
2. Spin-one Haldane gap
3. Parity is not in QAC^0
4. Constant-error Aaronson-Kuperberg conjecture
5. Unitary VOAs generating conformal nets. Things are changing fast. I cannot even imagine what will happen over the next few months, let alone a year. will depue: i asked GPT 6 Pro and Fable 5.1 to rank all discoveries in the last three years for Human discovered
for AI discovered (before October 6th)
for AI discovered from OpenAI/math repo 81% of them have been released today. wtf

Fable 5.1’s top 100 were 59% today’s list, 87% AI, full list at the link.

Sauers: My agents (Opus 5.5, Astra, and 5.6 Sol) already solved some of these beforehand (on GitHub). I wonder what % of these actually require their internal model?

The World Does Not Understand

The AIs all say when asked about all these math proofs as a hypothetical that this would be a huge deal, and should be front page news, historic beyond any reasonable comparison, the biggest day in mathematics. Some think it is impossible.

It turns out this was not front page news. Most people did not hear about it. It should have been, but no one in the news business knows what it means, or thinks people would care.

Kevin A. Bryan: Ok, finished going through the OAI math list. It is bonkers and should be front page news around the world if folks understood what this meant. But if I am running strategy at a lab, #1 priority is “do this for medicine, oncology, battery efficiency, etc as fast as possible”.

Joshua Gans starts off his coverage with the actual front page, in order to show us what things The New York Times thought were more important than solving 90 of the top 500 open math problems at once.

Seth Burn: Sometimes front page news doesn’t initially reach the front page. The best example is Sputnik. It launched 10/4/57, but the Soviets didn’t think it was a huge deal [and it got a few paragraphs of a right-side column]. It was only after America freaked out that it [became a banner headline] in Russia on 10/6/57.

It was a remarkably slow news day otherwise. Much of this is not exactly breaking. Yet no math is to be seen, even in the summaries at the bottom, which even include a generic AI thinkpiece.

Joshua Gans: What’s missing is this: 722 mathematics papers written by OpenAI. It isn’t an overstatement to say that this is probably the biggest day of scientific advancement in history. It is hard to evaluate, but the discussion that I have been seeing is that many of these are among the hardest and most significant results in mathematics. I suspect October 6th, 2026, will go down as some form of Judgment Day for AI in mathematics, but it portends so much more.

The counterargument is that there might be a lot of days like this:

roon (OpenAI): will yesterday be remembered as an important day? hard to say. in a punctuated exponential every local maximum looks invisible from a bit further out

For Now You Can Still Do Math

I like this model of why mathematics is not done quite yet:

Joshua Gans: The mathematicians are now going through this. To which I offer again the above Bookend slide to help:

We cannot yet show that the AI is better at conjectures or final verification. That will take some additional time, although probably not all that much time.

Verification has to be pretty good or we would have found a lot more errors by now. There are a lot of mathematicians who would love to find a mistake. We know, via the reasoning traces, that a big chunk of effort is going to verification.

You Will Need To Find A New Problem

Scott Aaronson reports on the Mathocalypse.

Scott Aaronson (Shtetl-Optimized): Last night my 9-year-old son was taunting my wife, complexity theorist Dana Moshkovitz, as follows: “mommy, I heard you got cooked! I heard that a robot solved the math problem you worked on for your whole career! OOF!” While my son was being a brat, he also wasn’t wrong. Whether you’re thrilled, depressed, angry, or whatever else about it, yesterday was surely one of the biggest days in mathematical history. And yes, among the 372 huge results released yesterday by OpenAI, on the recommendation of its advisory group of Timothy Gowers, Edward Witten, and other distinguished mathematicians, was a proof of Subhash Khot’s Unique Games Conjecture (UGC), a statement that my wife has worked toward proving for the entire time I’ve known her.

Then there is the other kind of ‘new problem,’ where those in denial, who three years ago took comfort in talking about how LLMs could not do math, have to find a new way to explain why anything an AI can do is not real or not meaningful.

Shtetl-Optimized: Experience has shown that, even now, there will still be people explaining in patronizing tones why none of this is real and none of it counts. If such people were capable of being impressed by anything that happens in the empirical world, of updating on anything, they would’ve already been impressed and already updated several years ago, long before things had reached the point of an actual Mathocalypse. So, they’ll say, maybe the alleged solutions are not solutions at all, but just “AI slop.” Or maybe none of the 372 well-known open problems that were solved were real math problems, they were all just glorified contest puzzles and trivialities. (After all, there’s still no Riemann Hypothesis!) Or maybe the entire 4000-year-old discipline of mathematics needs to be jettisoned: turns out that it was all just puzzle-solving and trivialities; all that’s different is that now the triviality stands unmasked. In any case, what really matters is that the true inner sanctum of human creativity hasn’t been breached and probably never will be, and also, that Sam Altman and Dario Amodei are contemptible little nerds. If you’re still a proponent of that doomed worldview, still aboard the sinking ship, I encourage you in the strongest possible terms to read yesterday’s other great contribution to AI discourse, besides the OpenAI Mathocalypse dump: namely, Scott Alexander’s open letter to Steven Pinker.

I do have to tip my hat to the fully general counterargument to all possible AI hype:

Cas: I think that AI smashing human achievements in mathematics might be overhyped, more an example of Moravec’s paradox than a robust sign of singularity soon. Recall how AI eclipsed humans at Go ~10 years ago but was much slower to get good at tasks that would change the world.

Whatever AI is good at is Moravec’s paradox. Whatever AI is bad at is why it sucks. Why are you so excited that AI is suddenly superhuman in an increasingly large set of domains, including the exact things we were saying it was dumb for being bad at? Surely this will not extend soon to other domains.

Verification or Evaluation Is Not Always Easier Than Generation

The steelman of this critique is to draw a distinction between verifiable versus unverifiable domains.

This argument says that whenever you have a verified domain, with known ground truth, AI will quickly become superhuman.

Whereas, if verification is difficult, or all you can do is evaluate in an informal way that requires a human in the loop to avoid biased errors and distorted behaviors, then AI capabilities will lag behind.

We were in a weird place for a few years, when pretraining dominated and LLMs were mostly trained on human words. This led to a seemingly non-Yudkowsky world in key ways, with LLMs being relatively excellent at a variety of useful but unverified domains, while being hopeless at math.

Now, with post-training dominating, LLMs are again stronger at math and coding, and things again look like you would expect. AI is quickly getting better at everything, but progress in other domains, while super fast compared to almost any other tech ever, is relatively slow for now. Thus, we advance math and coding and similar domains first, which then automate AI R&D, which then accelerates everything else.

Cracking the Code

Justin Drake notices that there is a striking under-representation of cryptographic breakthroughs among the results. As in, zero of the 719 abstracts mention cryptography, LWE or discrete logs.

One option is that AI is relatively weak at cryptography, another is luck or that OpenAI didn’t have cryptography in its initial problem set (perhaps to avoid exactly this issue?), or perhaps the government or OpenAI are censoring those results.

Justin calls for a ‘bunker mode’ for the blockchain industry, to protect against sudden mass breakthroughs.

Vitalik says not to panic, but that there are serious risks to cryptography from potential new math discoveries, including to lattices but not yet to hashes.

Vitalik Buterin: But for anything that has structure, you should assume that AI will make at least some progress in breaking that structure. Here, one reasonable inference is that if you want to make something plausibly long-term secure, multiply the key sizes by 10. … Theoretically, of course it’s possible that hashes are broken too (eg. P = NP would imply that). But I think P = NP is very unlikely. And intuitively, it’s much more likely that a mathematical object has exactly no exploitable structure (like hashes are intended to), than that a mathematical object has exactly ~3 forms of exploitable structure (for elliptic curves: associativity, Schoof, pairings) and not some secret fourth form of structure we have not yet discovered that greatly degrades its security (for elliptic curves, ECDLP and pairing security). Similar for LWE, SVP, RLWE and the zoo of lattice problems. For this reason, we do not yet see any reason to worry and start padding the byte size of hashes (if we start to worry more, we would pad the round count first before doing anything to the byte size). Concrete TLDR, my own personal views: * Hash-based > lattice-based, in those situations where hash-based is possible at all
* For anything lattice-based, be much more paranoid on param sizes. Remember that blockchains are only a small portion of the cryptography story; this point goes far beyond blockchains and applies to eg. access to websites, secure messaging, Tor / VPNs …
* For privacy protocols, strongly favor NOT putting encrypted notes onchain. Instead, send them offchain through some third-party mechanism.
* If it’s not difficult for you, keeping your funds in addresses which have not yet been used to make a transaction is a good idea. If it’s easy for you, do it. **But be careful about migrations; I personally have lost more money in botched migrations than I have lost in all hacks combined**.
* For multisig wallets, doing confirmations offchain is better than onchain, because this way the signatures of signer wallets do not get exposed to the public, so if ECDSA falls to AI much faster than expected, at least the multisig “gracefully degrades” to a 1-of-1 where the 1 is whoever was gathering the signatures – a much better place to be than “anyone can take the money”

I am guessing Vitalik’s point about botched migrations is highly underappreciated. It is very easy, in crypto, to lose fantastical amounts of money from stupid mistakes, the same way you can lose it from hacks or tech failures.

Yes, you should worry, especially given that OpenAI’s breakthroughs here did not involve trying maximally hard.

Kevin Madura: given the math announcements how many novel attacks on cryptography do we think exist now but are unreleased? What does this look like in 1 year? 5 years? I’d be eyeing DoD / NSA / NIST guidance pretty closely now Matthew Green: I think we might lose public key cryptography. Well, encryption specifically.

Never Change

Gary Marcus is defending his position by saying that the neural networks involved had a separate symbolic system, so none of this counts and he was right all along, and if ‘that’s over your head’ you probably shouldn’t be commenting on AI. The polite response is that the distinction does not matter in practice.

The Mathematicians Are Not Okay

Here are 100+ reactions from various different people in mathematics. A lot of them are very not happy about how OpenAI handled this, especially that so many of the papers were ‘unreadable slop’ rather than having been made nice first, or that OpenAI solved these problems at all. If you want to know ‘what are the mathematicians thinking’ this is a great resource.

Here is Terence Tao’s serious response, which focuses on how this disrupts the work:

Terence Tao: My feelings on recent developments are very mixed and complex. On the one hand, many of the AI-generated proofs appear to introduce clever new ideas that will be fruitful once digested, while also building upon the existing contributions of countless human mathematicians past and present. But at the same time, I am deeply frustrated that, in sharp contrast to traditional breakthroughs, none of the humans involved in these proofs are available to take questions, give talks, attend conferences, submit papers to journals, train students, or otherwise participate in the subsequent development of these results. Similarly, I am excited by the possibility of the community being able to use these tools to tackle ambitious and large-scale projects that one could not have even dreamed of in the past. But I am horrified by the many person-years of ongoing patient and deliberately slow research efforts – particularly by graduate students and postdocs – towards many motivating problems in mathematics being casually disrupted or destroyed by such a release. Much as one cannot unhear a movie spoiler or a crossword clue, one cannot explore a problem as profitably and richly once one is aware of an existing solution. Yes, one can still analyze and digest such an answer; but the best opportunity to do so is at the moment of its discovery, and such moments are increasingly wasted when delegated entirely to AI tools. And I mourn the path not taken, and the opportunities lost in the frantic race to develop this technology. Labs submitting their frontier models to independent researchers for proper scientific evaluation. Coordination with the research community to ensure these tools are applied to complement and enhance the abilities and activities of human researchers, rather than compete with them. Use of these tools to foster collaboration and sharing, rather than competition and secrecy. Opening new doors, without closing old ones. But that is not the path we now find ourselves in. Instead, the community needs to come together more than ever. To clearly declare our own standards and values, to build our own tools and practices, to support our most vulnerable members, and to chart our own path forward. Let’s get to work.

Here is a less polite version:

Here is another report from mathematician-land. People do not seem thrilled.

β/σi: > worried about their future
> worried about the solutions that **haven’t** been published (hello cryptography)
> in disbelief about the capabilities of these models
> sad the parts they fell in love with are gone (like we saw in swe)
> go solve “real” problems in bio (hilarious coming from mathematicians btw) ps this is just my small sample, i only have so many mathematician friends

Mathematicians seem to think of math as their playground, and rather than be happy about the problems being solved they feel their toys are being taken away. There is a struggle between wanting to know, and wanting to try to solve, and wanting the credit, and wanting solutions to be found at all. What do you actually care about most?

To be smart enough to be a mathematician, and also choose to be a mathematician instead of the many other better-paying things you could do, at least kind of requires some of the attitude that leads to cartoons where someone tells the mathematician someone found an application for their work, and the mathematician panics. Definitely #NotAllMathematicians, of course.

I am highly sympathetic. I really am. There could I have gone. That attitude has great value, not only inherently, but also exactly because following such curiosity leads to some of the most valuable discoveries that you would otherwise miss.

Will Kinney: One thing I’m really starting to notice is that the theoretical physics community is responding to the disruption created by AI in a completely different way than the mathematics community has. steve hsu: If your life’s work is a truly important problem – like curing cancer or fusion energy or discovering the true nature of quantum reality – then suddenly getting a solution from AI would be cause for great celebration. You might ponder for a moment the second order impact on your profession, but that would be overwhelmed by JOY for the gift that you and the rest of humanity have received.

It is also totally reasonable to focus on how things impact you and yours. That’s what hits home, and it is the part you can most control and are forced to face.

Josh Frisch: Like many other mathematicians, my main emotion thus far in 2026 has been loss: loss of meaning, loss of purpose, loss of the era of human proofs, loss of the ability to picture the future. With the release yesterday, October 6, 2026, of hundreds of beautiful results—many, maybe most, answering someone’s “one question”—I am trying to move beyond loss. There are so many beautiful results here: problems nobody had any approaches for, algorithms nobody thought could possibly exist, unexpected isomorphisms, constructions and proofs.​ Andrew Curran: Please resist the urge to mock the mathematicians struggling through this moment. Whatever your field, your expertise, your passion, you too shall one day experience what they are going through. And when you remember your words, you will feel remorse.

If your response is ‘well everyone’s passions are going away and the AIs come for us all’ then that’s worse. You know why that’s worse, right? Nor does it make this easier. There are some people out there reacting to these understandable reactions in truly vile ways. I urge them to stop.

Mathematics is facing a real problem here. If the AIs prove all the theorems, then our current methods of getting good at understanding math stop working. The traditional way you understand problems is by working to solve them and a solution is often not worth so much if no one understands it:

Jordana Cepelewicz (Quanta Magazine): It’s as if you were teleported to the peak of a tall mountain. Surrounded by fog, you have no idea where you are, or what’s around you. You do not know how your mountain connects to others, and you have no equipment to help you explore, no way to help someone else join you. If you had climbed the mountain yourself, you would have experienced how the human body adapts to altitude and changes in oxygen levels. You might have had to invent tools to navigate, to climb steep cliffs, or to make a shelter. You might have encountered a fellow explorer, gotten lost together in a hidden valley, and found a plant that could be turned into a life-saving medicine. Instead you’re perched on the peak but in the dark, while the maker of the teleportation machine tells you that it can explore the wilderness better than any human.

There are other ways to train understanding, but it could take a while to develop them, and they will require more motivation, especially intrinsic motivation.

This also destroys the system of PhDs and postdocs, since the problem you are looking to solve or the grant you want can get pulled out from under you at any time.

This also is a not fun place to be:

Many mathematicians have stopped posting open conjectures and potential ideas at the end of their papers, for fear that they’ll be scraped by bots and fed into AI models. Others are posting their papers before they’re ready in order to avoid getting scooped.

The Situation Turns Ugly

Some solutions were elegant and pretty cool. For example, the 9/4 matrix multiplication paper is 13 pages and bounds Strassen’s spectral characters directly. We’ve seen a bunch of people, after seeing their favorite problem get solved, share their ‘aha’ moment reading the proof.

Other solutions were less elegant.

I am with Wendigo that it is kind of weird and terrible that we are seeing all these new tiny improvements to previously elegant bounds on calculations, that have almost zero practical benefit.

Acer: oh yeah btw guys we can do integer multiplication faster than n log n lol I was definitely very surprised when this one came in lolWendigo: Am I the only one a little freaked out by the fact that our neat little lower bounds, nlog(n) in this case, are just slightly wrong. The sublime beauty of math is looking like a weak approximation resulting from our limited human faculties. Uubzu v4: I think the AI is just fucking with us when it takes an elegant bound like n(lg n) and improves it by 1/2^182. Like thanks for nothing buddy.

Now how about if we push it down to (1/2)^59? Then to (1/2)^34? And then (1/2)^14?

Anyone can keep doing things like this. All you need is a Codex subscription.

Doug Colkitt: We are publishing an update to OpenAI problem #109 (integer multiplication) with further tightening. κ = 2⁻³⁴ (tightened from κ = 2⁻¹8²) The exact witness is 8.3 × 10⁻¹¹, a roughly 48 million fold improvement over the previous result and a 2¹⁴⁸ fold improvement over original OAI result. The improvement came from removing the spacing penalty behind the quadratic bottleneck. This was done by moving compact control bits instead of entire windows. Beni: Doug when the fuck do you sleep Doug Colkitt: When my codex usage hits 0% of course. Konsti Wohlwend: Within a day of my prediction, the exponent went from: 0.𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟿𝟾𝟺 to 0.𝟿𝟿𝟿𝟿𝟺 “Weeks to months” was too conservative!

The Advisory Group Responds

Here is the full statement by the leading advisory group:

agmai.org: As announced a few weeks ago, OpenAI has released a large collection of mathematical results generated by an internal model, reporting solutions to hundreds of open questions. This is an important event for mathematics, with consequences both for mathematics and for the mathematical community that extend far beyond the individual results. AGMAI’s advisory role should not be interpreted as a judgment of the impact of these results or an endorsement of the process by which OpenAI obtained them. We do not speak on behalf of the entire mathematical community, and only the mathematical community can undertake the assessment that is needed. Making this work public is a first step. This release is the beginning, not the completion, of the process of human understanding and the incorporation of the work into mathematical knowledge. At the same time, the future of mathematical research cannot consist only of understanding results produced by AI labs. Mathematicians must be able to formulate their own questions, develop their own approaches, and explore directions that have not been selected as examples of an AI system’s capabilities. Equitable access to powerful research tools and adequate computational resources are essential to that freedom. We reaffirm our published recommendations on responsible release. We have discussed them with OpenAI and appreciate the company’s willingness to engage. While we consider these discussions constructive, it is ultimately up to the mathematical community to assess the extent to which our recommendations were followed successfully, and whether there are others we should suggest. We remain committed to engaging with any frontier AI lab on these questions and have already been in contact with several of them.

Another Mathematical Group Responds

Some look at what is happening, and cry out ‘NO!’

They are going to need a much better plan than this statement.

Do not be fooled into thinking this group, the ‘Association for Human Mathematics,’ is a big deal. There are hundreds of thousands of mathematics PhDs, and this group has 809 members. I include it because it perfectly embodies an attitude, nothing more.

Ariel: “Mathematicians did not ask for this work to be done” might become one of the historical quotes from this era, perfectly depicting the downfall of academia. Ethan Mollick: This document is going to be an assigned reading in college classes that cover this moment in time, there’s a lot happening in a few paragraphs… Association for Human Mathematicians: AHM Statement on OpenAI’s October 6 Release of Mathematical Documents Yesterday, on October 6th, 2026, OpenAI – which is currently defending lawsuits against accusations of illegal plagiarism, copyright infringement, and trademark dilution – released a repository of manuscripts purporting to contain solutions to a number of high-profile problems in mathematics. Mathematicians did not ask for this work to be done. The Advisory Group on Mathematics and Artificial Intelligence, from whom OpenAI has claimed to derive its legitimacy, opened their initial advisory statement by saying that frontier AI corporations should not test advanced mathematical problems on internal models. In ignoring the central premise of the Advisory Group’s position, OpenAI has indicated total disregard for the norms of scientific research — norms that guarantee that mathematics remains trustworthy, ethically researched, and in the public interest. Mathematicians have a particular vision of progress that is informed by history and field-specific considerations. We reject OpenAI’s assertion that this release advances our subject, and we urge mathematicians and the public to view the value of this publication model with due skepticism. Releasing over 700 files at once is not a demonstration of scholarship, but a demonstration of power. We urge mathematicians to discontinue their work with OpenAI and to return to a vision of science that centers human understanding. Association for Human Mathematics
Communications Working Group

Whatever you think of the rest of the statement, they are not wrong that the entire exercise here was directly against what the mathematics advisory group (AGMAI) wanted. AGMAI’s central ask was that OpenAI was asked to not test hard problems on private internal models. That request, regardless of what you think of it, got fully ignored.

A Different Approach

OpenAI admits that ‘dump it all on GitHub’ is not exactly meeting the advisory committee’s guidelines. They are exploring better options, and are hoping in the future to also have better-written papers.

On top of the big drop there was also this, that came out one day earlier, where Anthropic gave their solution to mathematicians to write up before releasing it, attempting to follow the advisory group’s recommendations:

New math result from Anthropic, then written up by Josh Alman and Virginia Vassilevska Williams, that is kind of a big deal. Anthropic had Claude research open problems, found the result, then hired Alman and Vassilevska Williams to write the paper as per what the mathematician council asked AI companies to do. Abstract: We give the first polynomial improvements over the textbook algorithms for 3SUM and All-Pairs Shortest Paths (APSP): we show how to deterministically solve 3SUM on n integers of polynomial size in O(n1.9992) time and APSP on directed n-vertex graphs with polynomially bounded integer weights in O(n2.9995) time. This refutes the 3SUM and APSP hypotheses. Using known reductions, we also refute the real-valued versions of the 3SUM and APSP hypotheses, the Exact Triangle hypothesis, the Zero-Weight k-Clique hypotheses, and the three rectangular hinted Online Matrix–Vector conjectures of van den Brand, Nanongkai, and Saranurak, and we give polynomial speedups for a variety of other problems. 𝖬𝖺𝗁𝖽𝗂 𝖢𝗁: The what now?! Atoosa Kasirzadeh: Could you write a tweet elaborating on what this could mean from your pov? 𝖬𝖺𝗁𝖽𝗂 𝖢𝗁: In a nutshell, suppose for a whole bunch of animals we knew that if they could whistle, then pigs could fly. But now someone saw a flying pig. That means we can now infer nothing about those animals (also in a sense, “dynamic programming isn’t optimal”). Scott Kominers: First off: . Utterly unreal result. Second: I’m seeing a lot of complaints about the fact that Anthropic contracted with human experts to write up the paper – but to a first-order, isn’t this what the “Responsible Release” guidelines circulated last week say to do? If Math wants to take the view that “AI labs have a responsibility to provide support, including funding, for the development of human understanding of the AI mathematical output that they release” then it’s going to have to become comfortable with these companies financing mathematicians to do just that.

There was also another stray new math result via GPT that came out right before the massive drop of other math.

Ethan Epperly: John Urschel just showed that the growth factor for Gaussian elimination with partial pivoting is ~n^{1/2} with high probability, resolving a long-standing conjecture of Nick Trefethen!

Three Withdraws

The three withdraws are all the consequences of one sign error.

I love the idea of a website called Retraction Watch. This is one place OpenAI executed well. Mistakes happen, and even if more are found – and there will almost certainly be more errors found as the unverified 58% is still largely unvetted – the error rate is impressively low. The important thing is to own the mistakes once they are found. Retractions are often a sign you are doing something right, not something wrong.

Alicia Gallegos (Retraction Watch): “I think the withdrawals and corrections will be viewed positively by the mathematics community, but it will take a lot more than that to earn back the trust they have lost,” Sutherland told us.

There have also been fourteen revisions.

Leaning Into Lean

I liked this explanation from Jon Stokes: Mathematics is great for advanced AIs, especially GPTs, because there is no ‘reward hacking.’ The answer is the answer, if you do something crazy and out of distribution but it works then it works.

Except, are you sure? One angle is that you can either find a proof or you can find a bug in Lean, and are you sure the math will always be the easier option when you face the world’s toughest open problems?

Jai: It’s just a next breakthrough predictor. A stochastic genius. Nobel autocomplete. Michael Roe: My guess is that finding exploitable soundness bugs in Lean is substantially easier than proving quasi-Riemann hypothesis. Given which, we should expect some of these proofs to be Lean exploits. Still, I’m also prepared to believe they’re real. (as a computer security person, “AI can find exploitable bugs in Lean” is actually more alarming news than “AI proves the quasi-Riemann hypothesis”. Security bug apocalypse coming right up.)

The bigger risk is that Lean does not confirm that the stated theorem matches up to the famous open problem. Your terms might be different, in ways that are non-obvious at first glance.

So far, Lean is holding up, as are all the papers with Lean proofs modulo one corrected side claim, which is rather impressive. The system works. There are 3 proofs that have broken so far, and all of them are in the unformalized 58%.

Of the nine results discussed up top, five are checked in Lean (quasi-RH, matrix multiplication, π, Unique Games and Hadwiger), as in the Lean-checked 42%.

At some point, you will give the model a task that is substantially harder than some other way to solve ‘the problem’ of you assigning the task, whether or not the alternative path involves taking over the world. Then you will be the one that has a problem.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论