722 math papers from an AI, mostly one prompt each: can you trust them?
OpenAI just put 722 math manuscripts on GitHub, all written by a model it has not released yet. They are grouped into 372 families of problems, drawn from a pool of around 4,000 tasks. The typical price tag per result: about three hours of ChatGPT Pro level compute. And here is the detail to sit with for a second: almost all of it came from a single prompt to a single agent.
The headline items include a result on a famous geometry problem (Kakeya, in four dimensions) and a weaker, partial result in the neighbourhood of the Riemann hypothesis. Weaker as in: this is not a proof of the hypothesis, so please do not read it that way.
The Lean problem
Lean is a language where a computer checks every step of a proof, so nobody has to take anyone's word for it. About 560 of the 722 papers have that treatment. The other 160 or so do not, and OpenAI itself warns they may contain errors.
So is this a genuine breakthrough, or one very expensive flex?
Honestly, a bit of both. If these results hold up, research level output at this scale from single prompts is something new. But a group of advisers at the Institute for Advanced Study asked OpenAI for the model's name, the exact prompts and the compute spent per problem. OpenAI gave averages, and said it is not bound by those recommendations.
And averages hide the thing you most want to know: how many attempts went nowhere? We are looking at the 722 that made it out. We do not see the rest. You can be impressed and still keep the question open. Both are true at the same time.