OpenAI’s public mathematics repository warns that its 722 AI-generated mathematical manuscripts sit at different stages of verification. Not every manuscript has a Lean formalization, and some unformalized results could contain issues.
Released on October 6, 2026, the collection was produced by an unreleased internal model. It spans 372 related result families and includes manuscripts, source files, formal proof artifacts, and selected summaries of the model’s reasoning.
This is a public verification event, not a model launch. Mathematicians now have concrete arguments to examine beyond the company’s description of its system’s capabilities. How much of the output will survive checking, establish genuinely new results, and become mathematics that other researchers understand well enough to use remains unresolved.
722 Manuscripts Do Not Mean 722 Separate Discoveries
OpenAI defines a “family” as a group of related papers that can include a principal result, companion arguments, consequences, or alternative proofs. Several manuscripts may therefore belong to the same underlying mathematical development. Neither 722 manuscripts nor 372 families is an independently established count of solved open problems.
The collection provides several ways to inspect its claims. An overview describes the families, a manuscript map connects individual papers to supporting materials, and the preprints directory contains PDFs, source files, and manuscript-specific citation and build instructions. Separate Lean materials document the available formal proofs and their verification configurations.
Subjects listed for the released reasoning summaries include the irrationality exponent of π, correlations of multiplicative functions, diluted spin glasses, and the three-dimensional relativistic Vlasov–Maxwell system. These labels illustrate the collection’s breadth; they do not confirm that every associated claim is correct or resolves its broad research area.
With the papers public, experts can inspect theorem statements, assumptions, citations, and arguments. They can ask whether a claimed result matches the original problem and whether earlier literature already contains the relevant ideas. Answering those questions requires scrutiny of individual manuscripts and result families, not a count of files.
Lean Support Helps, but Its Scope Must Be Checked
Many manuscripts have accompanying Lean formalizations, according to OpenAI’s README. A formal proof artifact can provide a computer-checkable argument, reducing reliance on a reader’s interpretation of mathematical prose. Still, “has Lean support” is not a blanket certification of everything in a paper.
The repository’s formalization catalogue describes itself as a catalogue of papers with a formalized main result. That wording is narrower than saying every statement, explanation, and consequence in every associated manuscript has been formalized.
Reviewers need to establish what the formal theorem says, which assumptions it uses, and how it corresponds to the natural-language claim. A formally checked argument supports the proposition actually encoded. It does not, by itself, settle whether that proposition captures the intended research question, whether the result is original, or whether the paper gives appropriate credit to prior work.
In a collection containing principal results, companion papers, and alternative arguments, a proof artifact associated with one paper should not automatically be treated as covering all related manuscripts.
OpenAI acknowledges incomplete coverage. The README says it will add Lean formalizations as they become available, warns that some unformalized results could have issues, and commits to trying to fix those issues quickly.
Verification must therefore be assessed manuscript by manuscript. Some claims come with formal evidence that can be examined and checked; others still depend on scrutiny of the written argument. The catalogue does not justify declaring the entire collection verified.
Even a successfully checked theorem leaves researchers needing to explain its ideas, relate them to existing methods, and determine what further work they enable. Formalization and mathematical understanding serve different purposes.
The Generation Account Is Useful but Incomplete
The repository provides more process information than a simple announcement of claimed results. Its README says OpenAI expanded its evaluations to open research problems after performance on existing mathematical evaluations saturated. Some outputs also build on earlier results produced by the models, so the collection should not be interpreted as hundreds of wholly independent, one-shot solutions.
For the vast majority of results, OpenAI reports using the same procedure with an unreleased internal model. Across the evaluation, the model was posed approximately 4,000 problems. The resulting outputs were organized into families and manuscripts and filtered for an appropriate level of significance.
Those figures do not establish a clean success rate. The approximately 4,000 inputs are problems, while the published units are families and manuscripts. Grouping can include consequences or alternative proofs, and the release applies a significance threshold. Dividing either published total by 4,000 would mix different units and obscure the selection process.
Each result used, on average, compute equivalent to three hours of ChatGPT Pro thinking with the internal model, OpenAI says. This is the company’s compute comparison, not a measured promise that a publicly available ChatGPT product can reproduce these results in three hours. It is also not a dollar cost or a complete hardware specification.
The repository lists ten abridged reasoning summaries, offering a view into selected results. They are not a complete record of the reasoning behind all 372 families or the evaluation’s unsuccessful attempts.
Disclosed exceptions to the main procedure include work on a zero-free region for the Riemann zeta function and work described as a proof of the Hodge Conjecture for CM abelian varieties. One zeta-function writeup was human-edited for readability, OpenAI says. The collection was therefore not produced through one entirely uniform process.
The model remains unnamed and unreleased. The README’s high-level account also falls short of a reproducible description of every attempt, failure, and selection decision. Outsiders can investigate the published output’s correctness without having equivalent access to the system that generated it.
AGMAI’s Objection Goes Beyond Finding Errors
The independent Advisory Group on Mathematics and Artificial Intelligence, or AGMAI, set out its concerns in recommendations published September 29, before this release. The group explicitly asks frontier AI labs to stop testing advanced mathematical problems on proprietary models inaccessible to the broader scientific community.
Its concern is an imbalance between production and understanding. A lab can generate mathematical arguments without the people prompting its system being able to understand, verify, or take responsibility for them. The wider community then faces the work of establishing correctness, explaining the arguments, checking attribution, and incorporating useful results into subsequent research.
AGMAI contrasts that situation with established mathematical practice. Authors are expected to understand their arguments, verify them, accept responsibility for the content, and explain significant developments through papers, seminars, and conference talks.
The group says it received more than 600 responses when seeking community feedback and that a clear plurality supported its recommendations. This demonstrates an organized body of concern, without establishing that all mathematicians share one position.
For AI-generated work that humans do not yet understand, AGMAI recommends detailed reporting of the model, prompts, summarized reasoning, time, and estimated compute cost. It also calls for formalization where possible, explicit statements of formalization status, and documentation of the problems attempted and how they were chosen.
Sources
- public mathematics repositorygithub.com
- formalization cataloguegithub.com
- recommendations published September 29agmai.org
- The Verge’s reporting on the releasetheverge.com





