OpenAI 用未发布模型在约88小时内给出 Navier-Stokes 存在与光滑性问题(七大千禧年难题之一)的解答,并通过 GPT-6 Astra 完成17小时 Lean
关于纳维–斯托克斯千禧年奖问题:https://openai.com/index/navier-stokes-solution/ (来源:https://news.ycombinator.com/item?id=49613262)OpenAI 的令人印象深刻的成果,他们使用一个未发布的模型提出了纳维–斯托克斯存在性和平滑性问题的解决方案:https://en.wikipedia.org/wiki/Navier–Stokes_existence_and_smoothness,这是七个千禧年奖问题之一:https://en.wikipedia.org/wiki/Millennium_Prize_Problems,自2000年5月24日起设有100万美元的奖励。
这一发现在一定程度上被Tristan Buckmaster的欺诈指控所掩盖。Tristan Buckmaster是纽约大学的数学教授,他曾与Levent Alpöge合作研究相关问题,Levent Alpöge是一位杰出的数学家,目前在Anthropic工作。
Tristan的投诉伴随着他们自己结果的仓促发布版本:https://mastodon.social/@tristanbuckmaster/117233413705701198。这里有一个PDF描述了事情的经过:https://cims.nyu.edu/~tristanb/statement.pdf。非常简短的版本是,Tristan和Levent几乎花了一年的时间研究这个问题,广泛使用了Claude和Codex(主要是GPT-5.6 Sol),然后在8月15日取得突破。数学界流言四起,Tristan和Levent听说OpenAI得知Anthropic解决了“一个重大未解问题”,于是他们联系了OpenAI,并了解到OpenAI有一个团队在研究相关问题,采用了类似的方法。引用Tristan的话:
我问了他们第一次发送提示的时间。OpenAI在一段时间内没有直接回答这个问题。最终达成的共识是,这个提示是在过去几天发送的,在我们的工作信息到达OpenAI之后。
我问模型是否以我们的Codex会话为训练数据或访问了我们的会话,我们在整个项目中将所有草稿都放入其中。我被告知模型不会查阅用户数据。我再次询问关于训练的问题,但没有得到回答。
事情从这里开始变得更加复杂。OpenAI团队提出可以等待Tristan发表,或者让他为他们的结果撰写论文,但明确表示由于OpenAI与他雇主的竞争关系,Levent将不会被邀请作为合著者。
以下是OpenAI对他们工作的描述:
在9月1日星期二,我们听到有传言称两个千禧年大奖问题已经被解决。受到这些传言和我们内部模型性能飞跃的启发,我们开展了一项工作,对其在所有未解决的千禧年大奖问题以及一些其他高影响力问题上进行评估。[...]
这些智能代理在9月5日星期六得出了它们的解决方案,大约在首批智能代理启动后88小时。通过GPT‑6 Astra进行Lean形式化和验证又花了额外的17小时。
在所有尝试的问题中,智能代理发送了490万条消息,使用了约3000亿个输出标记。在解决纳维–斯托克斯问题的过程中,智能代理发送了270万条消息,使用了大约1300亿个输出标记。
(我们不知道他们使用的内部模型的成本结构,但如果按GPT‑6 Astra的公开API价格计算,3000亿个输出标记将花费1,500万美元:https://www.llm-prices.com/#ot=300000000000&sel=gpt-6-astra。)
这里是他们对Tristan和Levent工作的看法(重点为原文所加):
我们的工作始于9月1日,在听到一个传言之后,我们后来意识到这与Anthropic员工Levent Alpöge和纽约大学数学教授Tristan Buckmaster有关。在我们完成整个项目和Lean验证(9月6日)之后,基于传言认为他们也有纳维–斯托克斯问题的解决方案,我们联系了他们,提出同时发布我们的结果并在联合公告中承认他们的优先权。[...]
在他们公开发布之前,我们(研究人员和智能代理)没有通过任何途径看到他们的工作——尤其是,没有访问任何特定用户数据来解决这个问题。虽然不太可能,但我们不能排除从他们使用我们产品中产生的去标识化数据可能有助于改进我们的模型:https://openai.com/policies/how-your-data-is-used-to-improve-model-performance/。然而,我们的证明有显著不同,即便在欧拉问题中,准确证明的结果也不同(强迫与非强迫)。
我对这里发生的事情的理解是,OpenAI 听说一些千禧年大奖问题已被使用大型语言模型(LLMs)解决,并将此视为展示其最新模型实力的机会,而没有过多考虑抢先一个已经使用 OpenAI 自己模型解决这个问题接近一年的团队的舆论影响。
这种情况似乎与目前计算机安全领域的状况相呼应。Anil Madhavapeddy 最近指出,如今只要有漏洞的传言就足以找到安全漏洞:https://anil.recoil.org/notes/rumour-is-the-exploit,因为如果有人知道某些软件存在未修补的漏洞,他们就可以让自己的代理去寻找它。数学领域现在也是这样吗?仅仅知道一个问题有未发表的解决方案,就可能引发数百万美元的 LLM 支出以抢先解决它。
这也凸显了我对这一切如何运作的持续困扰之一。当一个 AI 实验室说“我的数据被用来提升模型性能”时,那到底意味着什么?
我过去最喜欢的两个假设性问题是:
我现在对此最喜欢的假设性问题是:
这是 Simon Willison 发布的一篇链接文章,发布于 2026 年 9 月 8 日:/2026/Sep/8/。
每月赞助我 10 美元,即可获得一份经过策划的电子邮件摘要,内容涵盖本月最重要的 LLM 发展。
On the Navier–Stokes Millennium Prize Problem:https://openai.com/index/navier-stokes-solution/ (via:https://news.ycombinator.com/item?id=49613262) Impressive result from OpenAI, who used an unreleased model to produce a resolution to the Navier–Stokes existence and smoothness problem:https://en.wikipedia.org/wiki/Navier–Stokes_existence_and_smoothness, one of the seven Millennium Prize Problems:https://en.wikipedia.org/wiki/Millennium_Prize_Problems that have been subject to a $1,000,000 prize since May 24th, 2000.
The discovery is somewhat overshadowed by accusations of skulduggery from Tristan Buckmaster, an NYU mathematics professor who was collaborating on related problems with Levent Alpöge, an accomplished mathematician who currently works for Anthropic.
Tristan's complaint accompanied a hastily published version:https://mastodon.social/@tristanbuckmaster/117233413705701198 of their own results. Here's the PDF describing what happened:https://cims.nyu.edu/~tristanb/statement.pdf. The very short version is that Tristan and Levent worked on the problem for almost a year, making extensive use of Claude and Codex (mainly GPT-5.6 Sol), then had a breakthrough on August 15th. The mathematical rumour mill kicked into gear and Tristan and Levent heard that OpenAI had heard that Anthropic had resolved "a major open problem", so they reached out and learned that OpenAI had a team working on a related problem, with a similar approach. Quoting Tristan:
I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI.
I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.
It gets more complicated from there. The OpenAI team offered to wait for Tristan to publish, or to have him author a paper about their result, but were clear that Levent would not be invited as a co-author due to OpenAI's competitive relationship with his employer.
Here's how OpenAI described their work:
On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors and by the step change in performance of our internal model, we launched an effort to evaluate it on all open Millennium Prize problems and a few other high-impact problems. [...]
The agents arrived at their resolution on Saturday, September 5, about 88 hours after the first agents were launched. Lean formalization and verification took an additional 17 hours via GPT‑6 Astra.
Across all attempted problems, the agents sent 4.9 million messages and used about 300 billion output tokens. In the process of resolving the Navier–Stokes problem, the agents sent 2.7 million messages and used approximately 130 billion output tokens.
(We don't know the cost structure of the internal model they used, but 300 billion output tokens at public API prices for GPT-6 Astra would cost $15,000,000:https://www.llm-prices.com/#ot=300000000000&sel=gpt-6-astra.)
Here's where they provide their perspective on Tristan and Levent's work (emphasis mine):
Our effort began on September 1st after hearing a rumor which we later realized was related to Levent Alpöge, an Anthropic employee, and Tristan Buckmaster, a math professor at NYU. After the completion of our full project and Lean verification (on September 6th), believing from the rumor they also had a solution of Navier–Stokes, we reached out to them to offer a concurrent release of our result and to recognize their priority in a joint announcement. [...]
We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models:https://openai.com/policies/how-your-data-is-used-to-improve-model-performance/ . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).
My interpretation of what happened here is that OpenAI heard that some Millennium Prize problems had been solved using LLMs and saw this as an opportunity to demonstrate the power of their latest model, without thinking too hard about the optics of scooping a team who had been using OpenAI's own models to work on this problem for the best part of a year.
This situation appears to mirror what's happening in the world of computer security right now. Anil Madhavapeddy recently pointed out that Just a rumour of a bug is enough to find a security exploit these days:https://anil.recoil.org/notes/rumour-is-the-exploit, because if someone knows that some software has an unpatched vulnerability, they can set their agents the task of finding it. Is the same now true of mathematics? Just knowing that there is an unpublished solution to a problem might trigger millions of dollars in LLM spending to get there first.
This also highlights one of my ongoing frustrations about how all of this works. When an AI lab says that my data is "used to improve model performance", what does that actually mean ?
My two favourite hypothetical questions regarding this used to be:
My new preferred hypothetical for this is:
This is a link post by Simon Willison, posted on 8th September 2026:/2026/Sep/8/.
Sponsor me for $10/month and get a curated email digest of the month's most important LLM developments.