OpenAI published 722 AI-generated math manuscripts on October 6, crediting an internal model nobody outside the company can access. By the next morning, three papers were pulled. A sign error in one algebraic geometry proof had quietly cascaded into two dependent manuscripts, taking them down with it. Fourteen more papers were revised the same day. The math community had spent less than 24 hours with the release before cracks appeared.
What Broke and Why
The error originated in a paper titled “Algebraicity of Weil classes on split abelian eightfolds.” A flaw in the stabilization-trace cancellation argument invalidated the proof — and because two other papers depended on that construction, they collapsed with it. All three were retracted on October 7. The repository now lists 719 manuscripts, with 14 others carrying revision notes.
This matters beyond the specific papers. At the time of the retraction, only 42% of the 722 manuscripts had machine-checked Lean proofs. Lean verification confirms that a proof is logically consistent — it does not tell you whether the formalization matches the intended mathematical claim, whether the result is novel, or whether the reasoning is comprehensible. The other 58% had no formal verification at all. OpenAI’s own repository notes that unformalized results “could contain errors.”
The Verification Problem OpenAI Created
The deeper issue is not that an AI made a sign error. The issue is that you cannot check the work. OpenAI has not released the model that generated these results. Only 10 of the 722 papers include any reasoning summary. The rest are conclusions without an auditable process — you get the answer, but not the reasoning chain that produced it.
MIT’s Andrew Sutherland put it plainly: “Until and unless they release the model and people can replicate their results, I think you should treat any claims as unverified.” MIT’s Dor Minzer, whose own team rushed a 95-page paper in three days to avoid being scooped by OpenAI’s anticipated release, called the situation “math by press release.”
The Association for Human Mathematics (AHM) escalated further, urging mathematicians to stop working with OpenAI entirely and calling the release “total disregard for the norms that keep science trustworthy.” This followed a September open letter from 25 Fields Medalists warning that AI labs are undermining open research culture by racing to announce results without verification.
Not Everyone Agrees
The backlash is real, but not unanimous. Scott Aaronson called the release potentially “one of the biggest days in mathematical history.” University of Toronto’s Daniel Litt countered the boycott calls directly: “if answers exist, why withhold them from the community?”
The split reflects a genuine philosophical divide. One camp says AI-generated results without a reproducible process are epistemically worthless regardless of correctness. The other says verified conclusions are still conclusions, and mathematical progress should not wait for institutional comfort.
Why Developers Should Care
The math community is running into a problem that software is about to hit at scale.
A useful framing: this is a pull-request flood without code ownership. OpenAI generated roughly 4,000 proof attempts at about three hours of compute each, then dropped the successful subset on GitHub. Mathematicians now face months of expert review time to verify work that cost hours to produce. The producer/reviewer asymmetry is brutal — and it is the same asymmetry that emerges when AI generates more code than any team can meaningfully review.
Three things this episode makes clear for anyone building AI-assisted workflows:
- Reproducibility requires releasing the harness, not just the output. Shipping a binary without source code is not open source. Dropping 722 proofs without the model is not open science.
- Auditable AI output is non-negotiable. If you cannot retrace how a result was produced, you cannot fully trust it. The sign error here was caught because someone read the paper. The other 58% are still waiting.
- Verification bottlenecks are the actual constraint. Generation is cheap. Review is scarce. Any system that widens that gap degrades quality — in math and in code.
The Pattern
OpenAI is applying its product playbook to scientific research: move fast, publish, fix errors as they surface. That approach worked well enough for ChatGPT features. In mathematics — a field where a single sign error can invalidate a proof that 14 other papers depend on — the tolerance for that model is lower, and the cost of getting it wrong accrues to a community that did not ask for the load.
The remaining 719 papers may contain legitimate breakthroughs. The Unique Games Conjecture result alone, if it holds, would be significant. But until the model is released and results can be independently reproduced, the math community is being asked to carry the verification work that OpenAI did not do. That is not how scientific contribution is supposed to work. The feud between OpenAI and the mathematics community did not start here — but this week, it escalated.
Read the full release and repository at OpenAI’s math GitHub repository. Background on the broader AI-math tension is covered in depth by Quanta Magazine.













