AI at IMO 2025: a round-up
Setting the scene
The 2025 International Mathematics Olympiad has come and gone. Reminder: this is an exam for high-school kids across the world (each country typically sends six kids), comprising of two 4.5-hour exams each containing three questions, so six questions in total, which I’ll call P1 to P6. Solutions to each question are scored out of 7, for some reason lost in the mists of time. As you can see pretty clearly from the individual results and especially the sea of scores highly close to the sequence “7,7,7,7,7,0”, this year had 5 reasonable questions and one stinker, namely P6. Around half of the participants get medals, determined by scores chosen so that approx 1/12th participants get a gold, 2/12ths get a silver and 3/12ths get a bronze. I still have my gold medal from 1987 in a drawer upstairs; I beat Terry Tao (although I was 18 and he had just turned 12 at the time; he was noticeably younger than all other contestants). The cut-off for gold this year was 35, which is also the answer to 7+7+7+7+7+0. You can take a look at the questions for 2025 (and indeed for any year) at the official IMO website.
Back in 2024 Google DeepMind “took the IMO” in the sense that they waited until the questions were released publically and then throw two tools at them; AlphaGeometry, which is designed to solve geometry problems in a human-readable format, and AlphaProof, which is designed to write Lean code to solve any mathematics problem written in Lean format. Humans translated the questions into a form appropriate to each tool, and then between them the tools solved 4 out of the 6 problems over a period of several days (rather than the 9 hours allowed for the exam), which led to DeepMind reporting that their system had “ achieved Silver Medal standard” and which then led to the rest of the world saying that DeepMind had “got a Silver Medal” (although right now it is only possible for humans to get medals). Indeed DeepMind were one point off a gold in 2024.
This leads inevitably to two questions about IMO 2025. The first: will a system “get gold in 2025”? And the second: Were the IMO committee going to embrace this interest from the tech companies, define precisely what it would mean for an AI to “get gold”, demand an entrance fee or sponsorship from the tech companies and in return use official IMO markers to give official grades, also reporting on scores from tech companies who didn’t do very well? Spoiler alert: the answers are “yes” and “no” respectively.
Deciding the rules
Because the IMO committee chose not to write down any formal rules for how AI could enter the competition, this enabled each of the tech companies to make up their own rules and in particular to define what it means to “get gold”. In short, it enabled the tech companies to both set and mark their own homework. But before I go any further, it’s worth discussing what kind of entries are going to come in to the Wild West that is “AI at IMO 2025”. There are going to be two kinds of submissions — formal and informal. Let’s just break down what these mean.
“Informal” is just using a language model like ChatGPT — you give it the questions in English, you ask it for the answers in English, you then decide how many points to give it out of 7. Let’s scrutinise this last part a little more. IMO marking is done in a very precise fashion and the full marking scheme is not publically revealed; there are unpublished rules established by the coordinators (the markers) such as “lose a mark for not explicitly making a remark which deals with degenerate case X” and, because of this secrecy, it is not really possible for someone who is just “good at IMO problems” but who hasn’t seen the mark scheme to be able to accurately mark LLM output. Given that (spoiler alert) the systems are all going to get 0 points in problem 6, the claim that AI company X got a gold medal with an informal solution rests on things like “we paid some people to mark our solutions to questions 1 to 5 and in return they gave us 7/7 for each solution despite not having seen the official mark scheme”. Judging by this Reddit post there seems to have been some unofficial but failed(?) attempt by the tech companies to get the coordinators to officially give out scores to their informal solutions.
“Formal” is using an computer proof assistant, which is a programming language where code corresponds to mathematics. For such entries, someone has to translate the statement of each question into the language of the proof assistant, and then a language model trained to write code in this proof assistant will attempt to write a solution in code form. I had naively hoped that IMO 2025 would come with “official” translations of the questions into languages such as Lean (just as they supply official translations into many many human languages), but no such luck. So the formal solutions will involve something (probably a human) translating the question into the relevant computer language (possibly in many different ways; translation, just like translation between human languages, is not uniquely-determined by the input) and then giving the formal questions to a system which has been trained to write code solving them in the relevant language. Another catch here is that, as the name suggests, a proof assistant is something which can check proofs, and unfortunately the only question in the 2025 IMO of the form “Prove that…” was P2; all other questions were of the form “Determine this value”. This throws a spanner into the works of a proof-assistant-based attempt on the IMO: if one is asked to “determine the smallest real number c such that [something]” (which was what P3 asked us to do) then what is to stop a machine saying “the number to be determined is c, where c is the smallest real number such that [the same thing], and the proof that I am correct is that I am correct by definition”? It is actually rather difficult to formally even say what we mean by a “determine” question. Formal systems attempting the IMO thus typically use a second (typically informal) system to suggest answers and then get the formal tool to try and prove the suggested answer correct. Modulo checking that the corresponding “proof” question does correspond to the question being asked (and this does need checking), formal proofs are of a binary nature: if it compiles then it gets 7/7 because it is a computer-checked solution to the question which uses only the axioms of mathematics and their consequences. It is not really meaningful to ask for partial credit here, so anything other than a full solution gets 0/7 (although we’ll see an example of someone claiming to get 2/7 with a formal solution later).
The results are in!
First out of the starting blocks were the wonderful people at MathArena. These are people who don’t work for a tech company and are trying to do science. They reported on the performance of the best informal models available to mere mortals such as us (sometimes at a cost of hundreds of dollars a month), and the results were disappointing; the models that we regular people can get our hands on did not even get a bronze medal (this claim relies on the marking being accurate, which as I’ve already explained we cannot guarantee, but presumably they are in the right ball-park).
But of course we have not yet taken into account the fact that tech companies might have secret models up their sleeve, which have not been released to the public yet. And surprise surprise, this is what happened. So from this point on in the blog post, all claims made by tech companies are completely unable to be independently verified, a situation which is very far from my experience as a mathematician; one wonders if one is even allowed to call it science. Seems like Gregor Dolinar, the President of the IMO, is also cautious about the remainder of this post; his comments on AI entries can be seen here on the 2025 IMO website and I’ll reproduce them to save you the cli…