On 6 October, OpenAI put 722 mathematical manuscripts on GitHub: 372 results, from a model it has not released, that no journal has reviewed. I downloaded the catalogue and counted what can actually be checked.

OpenAI published more mathematics in one night than most journals publish in a year. The collection, at github.com/openai/math, is credited to an internal frontier model that is not named, not priced and not callable. The company says it is still being trained.

Some claims are enormous: a zero-free half-plane for the Riemann zeta function, which OpenAI brands the quasi-Riemann hypothesis; the Hodge conjecture for CM abelian varieties; the irrationality exponent of pi; both Mahler conjectures. If a fraction hold, this is a landmark. If not, it is the largest pile of unverified mathematics ever assembled in public.

Animated poster: a dark field of 722 small paper marks in a 38 by 19 grid. A stamp presses down and inks 162 of them in pale mint; the other 560 stay faint and blank. Counters read 722 published and 162 machine-checked.Animated poster: a dark field of 722 small paper marks in a 38 by 19 grid. A stamp presses down and inks 162 of them in pale mint; the other 560 stay faint and blank. Counters read 722 published and 162 machine-checked.
Watch: the stamp inks 162 of 722 marks. The rest stay blank. Design: KREO Journal.

What arrived, and how you are meant to check it

A number is not a proof, so the real question is verification. OpenAI is unusually forthcoming about the mechanism. It posed roughly 4,000 open problems, let the model think, and kept what cleared a significance bar. Each result used about three hours of ChatGPT Pro compute.

It also published Lean formalisations for part of the collection. The instructions to check them need three tools: comparator, lean4export, and landrun, a sandbox.

A number is not a proof. A pile of proofs is not yet a body of knowledge.

I counted the checkable part

I downloaded the catalogue and the formalisation manifest on 7 October. The repository's CONTENTS.md lists 722 manuscripts across 372 families, exactly as advertised. The lean/formalization.yaml, which catalogues papers with a formalised main result, lists 162 of them, with 185 declaration entries and 178 challenge files.

Manuscripts published722
Result families372
Papers with a formalised main result162
about one in five
From OpenAI's CONTENTS.md and lean/formalization.yaml, read 7 October 2026.

So 162 of 722, about one in five, carry a machine-checked core. The other 560 are words and PDFs. That is what a first release looks like, but the gap, not the headline, is the number that matters.

I then went to run the checker and stopped at the first gate. This Windows machine has no Lean and no lake, and landrun is a Linux sandbox. The command the README gives cannot run here. The tooling that would confirm any of this is out of reach for an ordinary reader.

Interactive · the gate

What it takes to check one result.

Three tools stand between the claim and a check. On the machine it wants, all three clear. On mine, the first jams.

  1. comparatordiffs proof against challenge
  2. landruna Linux sandbox, no Windows build
  3. lean4exportdumps the kernel proof

All three clear. The proof checks.Gate one jams: no sandbox, no Lean.

The honest counterpoint

None of this dismisses the work. A machine-checked proof is real evidence: if Lean accepts it under the three axioms OpenAI permits (propext, Quot.sound and classical choice), the logic holds. OpenAI is candid about the gaps. It says some unformalised results could have issues and that the Riemann and Hodge results fall outside its fixed procedure.

What none of it settles is significance. Lean checks that a proof is valid, not that the theorem matters or is new. The model is not released, so nobody can reproduce how it reasoned. Terence Tao warns that dumping results on the field may harm it; more than two dozen Fields medallists signed an open letter; Andreas Thom alleges dishonesty over attribution. MIT's Andrew Sutherland sets the bar: treat the claims as unverified until the model is out and the results are replicated.

What would change my mind: the model shipping, or an outside mathematician verifying one named result in public.

Readout: what is confirmed, and what is not
  • Confirmed (OpenAI, 6 Oct 2026): 722 manuscripts in 372 result families, published at github.com/openai/math under Apache-2.0, produced by an unreleased internal model, at about three hours of ChatGPT Pro compute per result.
  • Confirmed by my own count (7 Oct 2026): CONTENTS.md lists 722 manuscripts and 372 families; lean/formalization.yaml lists 162 papers with a formalised main result, 185 declaration entries, 178 challenge files.
  • Confirmed (OpenAI README): the Riemann zero-free region and the Hodge CM results are exceptions to OpenAI's fixed procedure, and the Riemann write-up was human edited. The repository has issues disabled and pull requests restricted to collaborators.
  • Reported (New Scientist, The Decoder, Engadget, QZ, 7 Oct 2026): mathematicians dispute the release; a Fields medallists' letter warns of harm to the field; Andreas Thom alleges dishonesty over attribution; MIT's Andrew Sutherland calls the claims unverified until replication.
  • Unconfirmed: every mathematical claim itself. No result is peer reviewed, and the model that produced them cannot be run by anyone outside OpenAI.
The whole verification contract
  • lean/ComparatorChallenges/QuasiRiemannHypothesis.json
    
    {
      "challenge_module": "ComparatorChallenges.QuasiRiemannHypothesis",
      "solution_module": "OAI.NumberTheory.DirichletL.Nonvanishing",
      "theorem_names": [
        "OAI.riemannZeta_ne_zero_of_seven_eighths_lt_re"
      ],
      "definition_names": [],
      "permitted_axioms": [
        "propext",
        "Quot.sound",
        "Classical.choice"
      ],
      "enable_nanoda": false
    }
  • This is the entire check for OpenAI's flagship claim: two module names, one theorem, three permitted axioms. Everything else is the model's word.

If you are working out what you can actually verify in your own stack of machine-learning tools, email brandon@kreostudio.co.uk.

Sources & references
  1. Sharing AI progress in mathematics, OpenAI, 6 October 2026.
  2. github.com/openai/math, README, CONTENTS.md, lean/formalization.yaml and lean/ComparatorChallenges/, read 7 October 2026.
  3. OpenAI announces 722 mathematical discoveries in one go, New Scientist, 7 October 2026.
  4. OpenAI dumps 372 AI-generated math proofs on GitHub, The Decoder, 7 October 2026.
  5. OpenAI just posted hundreds more results on major math problems, Engadget, 7 October 2026.
  6. OpenAI released 372 groups of math results on GitHub, Quartz, 7 October 2026.
  7. Recommendations, Advisory Group on Mathematics and AI, Institute for Advanced Study, 29 September 2026.

Reader signal

Was this useful?

Work with KREO Studio

AI engineering, data science and design architecture, from Plymouth to the wider UK.

Next Article

Open Weights · 5 min read

Can One Fat Cat Put Europe Back in the AI Race?