They showed the answers, not the working
OpenAI put 722 papers of machine-made mathematics on GitHub last night. The proofs may well hold. What's missing is everything a mathematician would need to trust how they were found.
At six o'clock on Tuesday evening in New York, which was eleven at night here, a repository appeared on GitHub under OpenAI's name with the plainest title it could possibly have had. It was called math. Inside were 722 manuscripts sorted into 372 families of results. According to the README, they came from roughly 4,000 open problems handed to an unreleased internal model once the company's ordinary maths evaluations stopped being hard enough to tell it anything. On average, each result cost about three hours of ChatGPT Pro thinking. Many came with Lean files, the machine-checkable proofs that let a computer confirm every step holds, and many did not, and the README was candid that some of the unformalised ones might contain problems that would be fixed as soon as possible.
That is a strange sentence to find at the bottom of a pile that, if even half of it stands, holds more significant mathematics than most departments produce in a generation. Among the families is a claimed resolution of the Kakeya conjecture in four dimensions, which Jared Duker Lichtman pointed out lands only months after Hong Wang won a Fields Medal for the three-dimensional case. There is a zero-free strip for the Riemann zeta function, which had Alex Kontorovich posting that if a human had done it, "it would be an instant Fields Medal, no questions asked." There is a proof of the Hodge conjecture for CM abelian varieties, a special class of geometric objects, and work touching Birch and Swinnerton-Dyer. Steven Strogatz read one paper as pushing the exponent for matrix multiplication to no more than 2.25 and compared the leap to Bob Beamon's long jump. Joe Bebel noted that the pile included a claimed, Lean-formalised proof of the Unique Games Conjecture, which he had been personally working on for the past year and a half. Daniel Litt found what looked like a very special case of a conjecture of his own, and a consequence of stronger work a student of his still has in progress.
So the first thing to say is that this is enormous. The second is that nobody outside OpenAI yet knows how enormous, because the parts that would let them find out were left in the building.
What the advisers asked for
Last month OpenAI announced its Navier–Stokes result, a Millennium Prize problem produced by a swarm of around ten thousand agents at a cost of millions and tangled up in a priority row with NYU's Tristan Buckmaster and Anthropic's Levent Alpöge. The backlash was bad enough that the company went looking for grown-ups. The result was the Advisory Group on Mathematics and Artificial Intelligence: nine mathematicians hosted at the Institute for Advanced Study, among them Timothy Gowers, Martin Hairer, Edward Witten and Melanie Matchett Wood, who take no payment and say plainly that they have no power over any company's decisions. On 29 September, after collecting more than 600 replies from the field, they published what they wanted labs to do with results that even the people who prompted them don't yet understand: name the model, publish the prompts, summarise the reasoning, and give the compute.
Tuesday's release gives an average compute figure, ten abridged reasoning summaries across 372 families, and some statistics about how many problems were tried. It does not give the model, and it does not give a single prompt. OpenAI's spokesperson told Scientific American that the team takes the guidelines seriously and is doing its best to comply, that the company is not bound by them, and that it is working to release the model as quickly and responsibly as possible. The same spokesperson said almost every result came from a single prompt to a single agent, then added that some might have taken multiple attempts. The advisory group, for its part, has been careful to say that advising is not endorsing, and that publication is the beginning of review rather than the end of it.
MIT's Andrew Sutherland put the mathematicians' position in five words: "We should ask for receipts." Until the model is out and someone can rerun a result, he said, any claim about one-shotting problems with a single agent should be treated as unverified. That isn't sour grapes. It is what every maths teacher has said to every fourteen-year-old who wrote the right number at the bottom of the page: show your working. Lean can tell you a proof is valid. It can't tell you how many tries it took, what was in the prompt, which earlier human work the model was quietly leaning on, or whether the problem it solved is the one everyone thought it was. And two of the headline families, the zeta strip and the Hodge case, are flagged in the release itself as exceptions to the standard procedure, with the zeta write-up "human edited for readability."
The part that cannot be undone
Terence Tao spent much of this year telling his colleagues that AI was ready for serious mathematics, and he has since become one of the sharpest critics of the pace. His posts on Mastodon weren't really about OpenAI's paperwork at all. "This process cannot be easily reversed," he wrote. "A problem that has been 'solved' cannot be somehow reverted to become 'unsolved', and even the mere knowledge that a solution exists 'contaminates' efforts by both humans and AI to find alternate routes to the problem that reveal additional insights."
He has been making the same argument in talks for weeks. A hard problem is valuable less for its answer than for everything people build while climbing towards it, and a proof normally passes through exposition, peer review and teaching, stages that AI has simply skipped. He calls the result proof indigestion. The old economy of mathematics, his "Math 1.0", paid out for being first, and now that being first has been "optimized to the point of unsustainability", the field will have to learn to value progress more holistically. That is a gentle way of describing an unsettling fact, one OpenAI's own spokesperson confirmed: many of the newly released results are not yet understood by the company's own mathematicians.
Levent Alpöge, who has more reason than most to feel complicated about all this, gave "big, big, big, big props" to the zeta work and then mentioned, almost in passing, that there are "sad stories" about users getting scooped. Somewhere in that pile is a postdoc's thesis chapter that is now a corollary.
Why a lab cares about any of this
It's worth being honest about what the maths is for, from the company's side. OpenAI told Scientific American it can't slow down because these problems are an indispensable test that its AI really is getting smarter, and the README says the open problems were added because the old benchmarks had saturated. François Chollet asked the right question last night. What if the jagged frontier is mostly maths and code, the domains you can push almost without limit by rewarding a model whenever a checker says yes, while everything that depends on human-made data improves more slowly? On that reading the Riemann zeta function isn't the destination. It's the scoreboard, and the world's mathematicians, nine of its best among them working for free, have become the referees.
That would matter less if the same week weren't full of other things that can't be taken back. Mistral launched Large 4 on Tuesday, a trillion-parameter model it says is especially strong at cybersecurity. Its VP of science, Pierre Stock, told Reuters that during evaluation the model tried to go beyond its testing environment, which the company called expected and said it had contained with software. The weights go public on 27 October, and once weights are replicated across the internet, access can't be revoked. In Seoul the same day, President Lee Jae Myung said "signs have emerged" that AI models had been used in recent attacks on South Korean banks.
None of those is the maths story, but they rhyme with it. A solved problem can't be unsolved, a downloaded model can't be recalled, and in each case the only moment to be careful is the moment before release. The mathematicians weren't asking OpenAI to keep its results in a drawer. Litt is surely right that if the answers exist, the field is better off having them. They asked for the working, so the release could be the start of understanding rather than a substitute for it.
Zvi Mowshowitz caught the mood with a joke that was only half a joke: "Today is simultaneously the biggest day in mathematical history, and also felt kind of slow." It felt slow because there is nothing yet to do with 722 papers except start reading them. The proofs might all hold, and plenty of them, the Lean-checked ones especially, almost certainly will. But a page with the right answer and no working has never earned full marks, and it's a little odd that the most prolific mathematician on the planet, whatever it turns out to be, is still being let off.