Wednesday, Oct 7 | --:--
Back to home

OpenAI Drops 722 Math Manuscripts from an Unreleased Model — 162 Lean-Checked

OpenAI published a GitHub repository of 722 manuscripts in 372 result families from an unreleased internal model. About 4,000 problems were posed; each result averaged roughly three hours of ChatGPT Pro thinking. The formalization catalogue lists 162 papers with a Lean-checked main result. The rest carry OpenAI's own warning that some unformalized work could have issues. AGMAI asked for model, prompt and compute per result; OpenAI is publishing averages and artefacts, not prompts, and says it is not bound by the recommendations.

Times of AI Desk 7 min read San Francisco, CA View as Markdown
Cover illustration for OpenAI Drops 722 Math Manuscripts from an Unreleased Model — 162 Lean-Checked

OpenAI just moved AI mathematics from showcase proofs to industrial volume — and left the field to referee hundreds of claimed results from a model nobody outside the company can run.

On October 6, OpenAI published openai/math: 722 manuscripts grouped into 372 families of related results, produced by an internal model it has not released. The README says the model was posed about 4,000 problems during the evaluation, and that each result used on average the equivalent of about three hours of ChatGPT Pro thinking. Work on a zero-free region for the Riemann zeta function and a proof of the Hodge conjecture for CM abelian varieties did not follow that fixed procedure; the zeta write-up was human-edited for readability. OpenAI's accompanying post says it consulted the IAS-hosted Advisory Group on Mathematics and AI (AGMAI), will fund workshops on AI-produced results, and is "working to responsibly release the model."

What the repo actually contains

  • 722 manuscripts / 372 families — OpenAI's framing is that results "resolve or make substantial progress" on open questions; AGMAI, via The Verge, characterises the release as solutions to "hundreds of open questions." Attribute both.
  • 162 papers with a formalized main result, counted from lean/formalization.yaml (catalog header: "papers with a formalized main result"; recount before reprinting — the repo says formalizations will keep landing).
  • 10 abridged reasoning summaries, including the irrationality exponent of π, Kaplansky's direct-finiteness conjecture in characteristic two, and the Mézard–Parisi formula.
  • An explicit README warning: "some of the unformalized results could have issues."

This sits on top of earlier OpenAI math drops: a Lean-checked Navier–Stokes blowup proposal and the ten Astra-era advances. Those were showcase scale. This is a catalogue.

Claims vs checks

Scientific American reports that claimed results include the four-dimensional Kakeya conjecture and progress toward the Riemann hypothesis. Those specifics are Scientific American's reading of the manuscripts; we did not re-inspect those PDFs. The Riemann work in the README is a zero-free region / progress, not a proof of the hypothesis.

An OpenAI spokesperson told Scientific American that almost every result came from a single prompt to a single agent, though some may have taken several attempts, and that OpenAI's own mathematicians do not yet understand many of them. That single-agent claim is a spokesperson statement with a caveat, not a README fact.

AGMAI recommended disclosing the model, exact prompt and compute behind each result. OpenAI is publishing average compute and some statistics, not prompts. The spokesperson said OpenAI is not bound by the recommendations.

Mathematicians quoted in Scientific American split: Andrew Sutherland (MIT) treats one-shot claims as unverified until the model is released; Daniel Litt (Toronto) called the release good for mathematics.

The trade

A single unreleased model produced several hundred claimed results at a reported average cost of about three hours of Pro-level compute each. That is a capability signal about whatever OpenAI ships next — and a stress test of how a field absorbs artefacts faster than it can referee them. About a fifth of the manuscripts (162 of 722) have a machine-checked main result. The rest carry OpenAI's own warning. "Solved hundreds of open problems" is OpenAI's and AGMAI's characterisation, not an independent verdict.

Limits

  • Model, per-result prompts and per-result compute are withheld.
  • Formalization count 162 is a point-in-time recount of lean/formalization.yaml; the catalogue is designed to grow.
  • Kakeya / Riemann manuscript specifics = Scientific American's reading, not our PDF audit.
  • 372 families are related-result groups, not a count of solved problems.

Sources

Prior Coverage

Earlier Times of AI reporting on this thread.

Scroll to continue reading