Published: October 9, 2026 | Reading Time: ~9 minutes | Channel: techminute
There is a particular kind of silence that follows a typo. The one you find after you've shipped, after everyone has read it, after it's been quoted. Now imagine that typo is a sign error — a single flipped plus — and the "ship" is 722 research manuscripts claiming progress on 372 of the hardest open problems in mathematics, published simultaneously by OpenAI on October 6, 2026, into a field that did not ask for them.
The silence lasted about a day.
On October 7, OpenAI updated the history file of its openai/math repository with a withdrawal notice. Three manuscripts — gone, with their READMEs now carrying notices explaining "the gap" and linking to archived versions. The proximate cause is almost comedically small: in a paper titled "Algebraicity of Weil classes on split abelian eightfolds," a sign error invalidated what the changelog calls a "stabilization-trace cancellation argument." And because two other papers built their constructions on top of that argument — "Algebraicity of Kuga–Satake Correspondences for K3 Surfaces" and "The rational Hodge conjecture for products of K3 surfaces" — the whole stack came down. Three papers, one minus sign.
If you want a single image for the state of AI-generated mathematics in October 2026, that's it: a machine that can produce, in hours, work that would crown a human career — tripped up by arithmetic a first-year undergraduate is trained to catch.
Let's rewind 24 hours, because the withdrawal only makes sense in light of what OpenAI did first.
On October 6, the company pushed a GitHub repository containing — depending on when you counted — 722 manuscripts organized into 372 "families" (related papers sharing a principal result, companions, consequences, alternative proofs). The current catalogue, post-withdrawal, lists 719 manuscripts. The claimed results span geometry, algebra, number theory, and theoretical computer science: the irrationality exponent of π, the symmetric and general Mahler conjectures, Kaplansky's direct-finiteness conjecture, the Mézard–Parisi formula for spin glasses, isomorphism of free group factors, quasipolynomial bounds for arithmetic progressions, partial progress on several remaining Clay Millennium Problems. The repo even ships abridged reasoning traces for ten families, so you can watch the model think — or at least, watch an abridged summary of it.
The production numbers in the README read like science fiction, and they're the scariest part. According to OpenAI, the "vast majority" of results came from one unreleased internal model, at an average cost of three hours of ChatGPT Pro-level thinking compute per result, from roughly 4,000 problems posed during the evaluation. Three hours. A mathematician's career, compressed into an afternoon. Scott Aaronson, the UT Austin complexity theorist, compiled his own figures and put the broader test set at about 8,000 problems with a roughly 5% success rate — and noted, citing unnamed sources, that cryptography is conspicuously absent from the public release even as AI companies quietly probe cryptographic protocols behind the scenes. (That last part is Aaronson's reporting, not confirmed — I'm labeling it as such.)
To be fair, the repo is explicit about what it is and isn't. Not everything carries a Lean formalization — the machine-checkable proof artifact that would make a result, well, proven in the sense a proof assistant can vouch for. "Some of the unformalized results could have issues," the README admits, in what turned out to be the understatement of the week. At last count, about 42% of top-line results are formalized — 300 of 719 — after six new formalizations were added in the October 7 update.
And no, despite the breathless headlines: there is no complete solution to any remaining Millennium Prize Problem in this batch. Partial progress only.

Now the withdrawal itself, because the details matter more than the drama.
The history.md entry is worth reading in full, because it's the most honest document OpenAI has published in weeks. Beyond the three withdrawals, the same update revised 14 other manuscripts — "proof repairs, corrected statements, clearer hypotheses and dependencies, and one correction to an obsolete citation" — and updated 13 more just to cite the revised editions. One sign error, three dead papers, fourteen wounded, thirteen collateral. In peer-reviewed mathematics, a correction cycle like that takes months or years. Here it happened before the first press cycle finished.
The withdrawals were announced on X by Dan Roberts, an OpenAI research lead, who wrote that the company will "continue to update the repo with new formalizations and with any errata we notice." A spokesperson told Retraction Watch the errors were found during an audit — and, crucially, revealed the process design: OpenAI's collaborators at the Advisory Group on Mathematics and Artificial Intelligence (AGMAI) had recommended releasing the results without waiting for "full formalization," and roughly 50% of the results were released unconfirmed. Read that again: half the dump was, by OpenAI's own account, unaudited at publish time.
Alex Townsend, a Cornell mathematician, told Retraction Watch the cascade didn't surprise him — he expects more errors — and offered the sharpest process critique I've seen: OpenAI should have announced the Lean-verified manuscripts first, as their own announcement, and released the rest separately with an explicit request for community help. Instead, verified and unverified arrived in one indistinguishable avalanche.
Credit where it's due, because this story is genuinely two-sided: Andrew Sutherland, MIT research scientist, called the fast withdrawal "the responsible thing to do." OpenAI preserved the withdrawn papers' archived versions rather than memory-holing them. The changelog is public. This is, mechanically, better retraction behavior than plenty of human institutions manage. And yet.
The mathematics community's answer arrived on October 7, and it was not a grateful thank-you note.
The Association for Human Mathematics — chaired by Fields Medalist Terence Tao — issued a statement urging mathematicians to discontinue work with OpenAI. Its most-quoted line deserves the quotation: "Releasing over 700 files at once is not a demonstration of scholarship, but a demonstration of power." The statement's opening also pointedly nods to OpenAI's ongoing copyright lawsuits — the implied question being whose life's work trained the model that just automated theirs.
Tao's own writing this week is the intellectual center of gravity. His "Math 1.0 → Math 2.0" framing holds that mathematics used to prize being first to solve an open problem, and that this goal has now been "optimized to the point of unsustainability." What AI users are doing, in his words, is harvesting — solving problems autonomously, en masse, by people with no interest in the field, leaving fewer seminars, fewer collaborations, fewer young researchers drawn in. And the damage is asymmetric: once a problem is believed solved, it can't be un-solved. Even knowing a solution exists contaminates the search for alternative approaches — the alternative approaches are often where the deep new mathematics lives. Twenty-five Fields Medalists, including Peter Scholze, Maryna Viazovska, and Martin Hairer, had already signed a statement warning of "severe misalignment" and raising "severe attribution and plagiarism questions" over rushed announcements.
Then there's the human toll, which Aaronson made concrete via his wife, Dana Moshkovitz — a theorist who has spent her entire career on the Unique Games Conjecture. The release includes a claimed proof of it. Her review, texted to Aaronson the night of the release: "It feels like something written by someone who's on psychedelics." The paper cites works without explaining why they apply, relies on what she called "some alien craziness," and is — her verdict — "so horribly written that it's impossible to read it without AI help." Sit with that: the claimed solution to a person's life problem arrived in a form that person cannot read without the help of the machine that wrote it.
Aaronson christened the whole situation "The Mathocalypse." And to his credit he doesn't pretend there's a clean side. He contrasts OpenAI's "dump it all" approach with the "Anthropic model" — where the company partnered with two algorithm researchers whose AI supplied the key idea for disproving two decades-old conjectures, compensated them, and co-wrote a version humans could follow. Anthropic's version has its own problem, Aaronson notes: a private company deciding which mathematicians get anointed as the human translators of machine results. Dump-and-run leaves the community doing free janitorial work on alien proofs. Partner-and-publish lets a corporation pick the emissaries. Pick your dystopia.
Sutherland's longer view, also to Retraction Watch: the withdrawals "will be viewed positively," but "it will take a lot more than that to earn back the trust they have lost" — trust burned, in large part, by OpenAI's rushed September 8 Navier–Stokes announcement, after which more than 8,000 researchers endorsed concerns about exactly this style of release.
Here's the thing I can't shake: the retraction took 24 hours, and that's the most hopeful fact in this entire story.
For most of the history of science, the bottleneck on knowledge was production — finding the proof took years, so errors hid for years. The openai/math repo inverts that. Production is now nearly free; three hours of compute per result. The bottleneck has moved entirely to verification, understanding, and trust — the slow, human, unfundable parts. OpenAI shipped 722 manuscripts and half of them were unconfirmed at publish. The verifier — Lean, the audit, eventually the community — is now the rate limiter for all of mathematics. The dump happened because it could; the withdrawal happened because, for once, the verification layer actually worked, and worked fast.
But notice what the 24-hour retraction did not fix. Three papers died; the other 716 are still sitting there, ~58% of them unformalized, in a repo with 12.8k stars, waiting for humans with careers and students and finite lives to sort through them. The sign error was the cheap error to catch — arithmetic is exactly what formalizers are good at. The expensive errors are the ones Moshkovitz flagged: citations that don't hold, constructions that don't mean what they appear to mean, proofs that are technically unassailable and completely unenlightening. Lean can verify that a proof is correct. Nothing can verify that it's good.
That's the real "Math 2.0" problem, and it won't stay in mathematics. Every field where production is being automated — code, drug discovery, legal analysis, this very blog post — is inheriting the same inversion: infinite plausible output, scarce trustworthy understanding, and a verification layer staffed by the same humans who were just told their specialty now takes three hours. The mathematics community's revolt isn't Luddism. It's a preview of every profession's negotiation with the dump.
OpenAI's spokesperson said, "We welcome scrutiny and feedback from the mathematical community." The community has now provided both, in industrial quantities. Whether anyone at OpenAI reads 591 Hacker News comments with the same diligence the model applied to 4,000 problems — that's the experiment running now.
A sign error is trivial. The question the sign error surfaced is not: who is going to read all this, and what does "solved" even mean if the answer is no one?
All claims verified against Gold-tier (the openai/math repository, its history changelog, and primary posts by the researchers involved) and Silver-tier (Retraction Watch, The Decoder) sources. Each source URL was scraped and confirmed accessible with full content on October 9, 2026. Aaronson's cryptography remarks are his own sourcing and are labeled as such; OpenAI's 4,000-problem figure and Aaronson's 8,000-problem figure are attributed separately, not averaged.