OpenAI says an unreleased internal model, running large numbers of coordinated agents, produced a resolution to the Navier–Stokes existence problem — one of the seven Millennium Prize Problems, million-dollar open math questions posed in 2000. Simon Willison’s commentary treats the result as less a math triumph than a case study in how frontier labs race, and what your data becomes when you work inside their tools.
The timeline is the story:
- On September 1, OpenAI heard rumors that two Millennium problems had been solved — rumors it later linked to a pair of human mathematicians, Tristan Buckmaster (NYU) and Levent Alpöge (Anthropic), who had worked on related problems for almost a year using Claude and Codex
- OpenAI launched agents on all open Millennium problems; they arrived at a Navier–Stokes resolution on September 5, about 88 hours later — roughly 2.7 million agent messages and 130 billion output tokens spent on this problem alone (about $15M at public API prices)
- Verification in Lean, a system that mechanically checks proofs, took 17 more hours
The controversy: Buckmaster and Alpöge had their own breakthrough on August 15 and allege OpenAI’s sprint began only after word of their work spread. OpenAI says its agents never saw their work but explicitly “cannot rule out that de-identified data derived from their usage of our products helped improve our models.” Alpöge was not invited as co-author — OpenAI cited its competitive relationship with his employer, Anthropic.
Willison’s read of the episode: OpenAI heard a rumor and saw a chance to demo its latest model, without thinking hard about the optics of scooping a team that had spent a year inside OpenAI’s own products. He draws a parallel to security research, where a mere rumor of a bug is now enough to trigger an automated hunt — is the same true of mathematics, where knowing an unpublished solution exists invites a multi-million-dollar race to get there first?
The deeper point is about trust in the toolchain: if you use ChatGPT to help partially solve a hard problem, what are the odds a later model helps someone else finish it first? Labs say your data is “used to improve model performance” — Willison’s recurring frustration — and this episode is the sharpest example yet of why that vagueness matters to anyone doing serious intellectual work with AI assistance.