OpenAI dumps 372 AI-generated math proofs on GitHub, telling the academic world to keep up

OpenAI has uploaded 372 mathematical results to a GitHub repository, claiming each one solves an open problem or makes substantial progress toward…

By Vane October 7, 2026 3 min read
OpenAI dumps 372 AI-generated math proofs on GitHub, telling the academic world to keep up

OpenAI has uploaded 372 mathematical results to a GitHub repository, claiming each one solves an open problem or makes substantial progress toward one.

The collection includes advances in computer algorithms and work related to the Riemann hypothesis. Most of these outputs came from a single prompt to a single AI agent, with each proof consuming about three hours of ChatGPT Pro Thinking compute on average.

The company is hosting the work in a GitHub repository complete with revision logs and citations. OpenAI states the same model already produced a solution to a Navier-Stokes problem that has been under formal review for weeks.

Formal verification could ease the review bottleneck

Many of the proofs include formalizations in Lean, a programming language built for machine-checkable mathematical proofs. More formalizations are planned because the volume of AI-generated results could easily overwhelm the math community’s capacity for manual review.

OpenAI also published details on its methodology, including summaries of the reasoning process, statistics on how many problems the model attempted, and estimates of compute costs.

Traditional journals aren’t built for this pace

OpenAI put its results on GitHub instead of peer-reviewed journals. It is a statement move, since it implies that the traditional process of doing science is too slow for this volume of potentially new knowledge.

OpenAI consulted with the Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study and loosely followed their public recommendations. The company did not publish any prompts and only shared average compute costs rather than per-problem figures. The advisory group includes Fields Medal winner Timothy Gowers among other renowned mathematicians.

The company also announced plans to fund workshops and conferences focused on understanding AI-produced results. OpenAI acknowledged it wants to improve the quality of its citations and presentation, and says it is working on a responsible release of the model to directly empower scientists with state-of-the-art capabilities.

OpenAI set one significant boundary for the advisory group beforehand, though: the mathematicians can advise on how results get communicated, but not on whether or how fast they are produced.

The math community is split on what mass-produced proofs actually mean

Lean formalizations can verify logical correctness, but they cannot judge whether a result is mathematically relevant or original. OpenAI is betting that its results push the boundary of human knowledge. Whether the mathematical community agrees remains an open question. So far, reactions range from excitement to frustration.

In a recent open letter titled “A Severe Misalignment of AI in Mathematics,” 25 Fields Medal winners warned of a deep disconnect between the AI industry’s goals and those of mathematics. Problem-solving, they wrote, is merely a tool and proxy for the real goal of conceptual understanding and insight. Mass-producing true statements could destroy fertile ground rather than bring new ideas to life, they argued. The effect would spill over into other fields.

Gowers has warned that within one to two decades, mathematical literature could grow enormously while no human community remains that truly understands it. Fields Medal winner Terence Tao has added that training young mathematicians needs to emphasise the human side and tightly limit AI tool use so that genuine learning and understanding survive.

What it means

For researchers, this changes the workflow from waiting months for journal publication to checking code repositories. However, it also means the burden of verification shifts entirely to the academic community, who must now decide which of these thousands of generated proofs are actually useful.

Scroll to Top