OpenAI released hundreds of machine-generated mathematics and theoretical computer science results in a GitHub repository on October 6, 2026. The new model produced hundreds of manuscripts after attempting roughly 4,000 problems, drawing both skepticism from independent researchers and new questions regarding model transparency and research pace.
The floodgates opened at 6 P.M. EDT when OpenAI revealed 372 research families comprising 722 manuscripts in a public GitHub repository. The collection spans pure mathematics, theoretical computer science, and mathematical physics, tackling problems that demand lengthy chains of logical reasoning.
OpenAI Releases 372 Research Families and Thousands of Evaluated Problems
During the research process, the model attempted roughly 4,000 problems, with an accepted result consuming about three hours of equivalent ChatGPT Pro thinking compute on average. OpenAI grouped related manuscripts into research families to illustrate how individual results connect, while also providing abbreviated summaries of the model’s reasoning for 10 distinct families.
The subjects covered in the repository range from number theory and complexity theory to geometry and mathematical physics. Specific examples include a result addressing the irrationality exponent of pi—which measures how closely rational numbers can approximate pi—along with contributions to NP-hardness, Mahler conjectures, arithmetic progressions, free group factors, quantum Heisenberg ferromagnets, and relativistic Vlasov-Maxwell equations.
Lean Verification Adds Code-Checking to Complex Proofs
Because automated generation can introduce errors, the repository incorporates an important verification layer using Lean, a programming language that validates a proof’s logic. Researchers can translate mathematical statements and proofs into formal code so a computer can check whether each logical step follows from the stated assumptions.

While many manuscripts still lack formal Lean versions, OpenAI says it will add more formalizations as verification work continues. The repository also tracks paper revisions and provides citation guidance to help researchers follow corrections or expansions of earlier manuscripts. The company developed its release process with advice from the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study.
Skepticism Persists Over Proprietary Models and Single-Agent Claims
Despite the technical verification offered by Lean, mathematicians have raised concerns about transparency. A spokesperson for OpenAI told Scientific American that an internal frontier model not released to the public produced almost every result in response to a single prompt handed to a single AI agent. That contrasts sharply with the company’s earlier Navier-Stokes solution, which required a 10,000-strong agentic swarm costing millions of dollars in computing power.
Sutherland added bluntly, We should ask for receipts.
Independent researchers have noted that the deluge will take months to parse, particularly in determining whether the proofs contain novel ideas or merely mash existing techniques. Terence Tao has vocally criticized frontier AI labs for the intense pace of their machine-generated results, while OpenAI acknowledged that many of its newly released findings are not yet fully understood by its own internal mathematicians.
Advisory Guidelines Clash With Selective Disclosure Practices
The release follows months of controversy surrounding OpenAI’s handling of automated mathematical discoveries. Following disputes over the Navier-Stokes existence and smoothness problem—which involves whether smooth, three-dimensional fluid flows can maintain smooth behavior or develop a singularity with infinite velocity—OpenAI announced an independent advisory group on September 21 to recommend responsible publishing practices.
Those recommendations state that companies releasing AI-generated mathematical results should publish the model, exact prompt, and compute time behind each finding. However, OpenAI chose to reveal only average compute times and additional statistics without sharing the underlying prompts, maintaining that while it takes the guidelines seriously and is working to release the model as quickly as possible, it is not bound by the advisory group’s recommendations.
Mathematicians Debate the Ethics of Withholding Versus Dumping Proofs
The rapid pace of AI output has forced the mathematical community to confront difficult questions about access and openness. Daniel Litt expressed support for immediate visibility, arguing that withholding the data serves no purpose.
Litt added, To me, it’s going to be a good thing for mathematics.
Meanwhile, OpenAI maintains that solving these math problems serves as an indispensable test to prove its AI systems are genuinely advancing in capability, signaling that the company plans to continue releasing results as it evaluates feedback from workshops and conferences.