Xena

RSS: https://xenaproject.wordpress.com/feed/
数学家们通过实践学习 Lean 定理证明器。

Anthropic抢先完成了费马大定理的Lean形式化证明

Anthropic用AI在11天内完成了费马大定理在Lean中的完整形式化证明,终结了Freek Wiedijk二十年前提出的100个形式化挑战。证明基于1995年Darmon-Diamond-Taylor的 exposition,代码超过1340万行,编译耗时是Lean数学库的20倍。 这项工作的数学价值有限——证明忠实跟随早期文献,没有带来新数学结论。但它的真正意义在于展示了autoformalization的能力边界:如果数千页文献能在11天内由AI完成端到端形式化,未来现代研究的形式化将实时发生。 一位EPSRC资助的费马大定理形式化研究者对此的反应是:数学上他99.9%确信证明无误,但AI自动形式化硬材料的能力将彻底改变数学论文的审查流程,并可能暴露Langlands纲领中某些"专家皆知"的假设是否真的成立。
评论点赞收藏25 天前

人类数学家正在被“反例”击败

<p>这几周反例挺有趣的。这篇文章基本上是我对形式化、人工智能工具,尤其是反例领域动态的看法。</p><p>单位距离</p><p>两个月前的今天(2026年5月20日),ChatGPT 在离散几何中推翻了埃尔德什的单位距离猜想。 这已经是老生常谈了,但我总得从某处开始。 该伴随着人类数学家的证词, 其中许多人我认识,也有少数我信任,他们相信该论点(他们曾提前接触并核实过)。 该证明的基本结构是,1960年代戈洛德和沙法列维奇提出的数论中一个深刻定理可以用来构造该猜想的反例。</p><p>距离我经历了中年危机、意识到自己在技术细节上不再信任许多人类数学家、发现精益, 并开始主张交互式定理证明器应在数学未来发挥重要作用,已经过去了9年。 所以我当然第一个问题是“反例在精益中是否被形式化”。 答案是“没有”。</p><p>但不到一周后(2026年5月26日),我收到了菲尔兹奖得主迈克·弗里德曼的邮件。 迈克现任首席科学官这家公司由图灵奖得主、 “人工智能教父”严乐村共同创办。Mike告诉我,他们的系统已经自动形式化了整个ChatGPT生成的论文在精益模式中, 我可以看看吗?我查了,我的博士后Thomas Browning也看了。这正是逻辑智能所做的:他们正好形式化了数论深奥定理蕴含埃尔德什反例的陈述。 突破性的大型语言模型生成数学正在实时形式化。 有趣的数据点。</p><p>当然,这里有一个大问题,那就是数论的深刻定理,它需要100+页才能证明(它需…</p>
评论点赞收藏71 天前

Formalizing Fermat workshop

I’m organizing a workshop in London on July 6th to 10th (2026) whose goal is to work on my EPSRC-funded project formalizing Fermat’s Last theorem in Lean. The initial aim of the project was to reduce ...
评论点赞收藏138 天前

Accelerating mathematics

Let’s say that someone had a big pot of money, and wanted to use it to accelerate mathematical discovery. How might they go about doing this? The traditional approach Historically it has been governme...
评论点赞收藏233 天前

AI at IMO 2025: a round-up

Setting the scene The 2025 International Mathematics Olympiad has come and gone. Reminder: this is an exam for high-school kids across the world (each country typically sends six kids), comprising of ...
评论点赞收藏423 天前

Think of a Number

My feed was recently clogged up with news articles reporting that Sam Altman thinks that AGI is here, or will be here next year, or whatever. I will refrain from giving even more air to this nonsense ...
评论点赞收藏472 天前

Think of a Number: An Update

A month or two ago I wrote this post which expressed my frustration with various issues around private datasets as a way of measuring the mathematical abilities of language models. More generally I wa...
评论点赞收藏562 天前

What is a quotient?

Undergraduate mathematicians usually have a hard time defining functions from quotients in Lean, because they have been taught a specific model for quotients in their classes, which is not the model t...
评论点赞收藏597 天前

Can AI do maths yet? Thoughts from a mathematician

So the big news this week is that o3, OpenAI’s new language model, got 25% on FrontierMath. Let’s start by explaining what this means. What is o3? What is FrontierMath? A language model, as probably m...
评论点赞收藏646 天前

登录芦苇

登录后关注作者、收藏内容和参与讨论。