OpenAI published 372 claimed mathematical results -- grouped from roughly 720 individual files across 16 areas of mathematics and theoretical computer science -- to a public GitHub repository on October 6. A company spokesperson said nearly all of them came from a single prompt handed to a single AI agent, running an internal model OpenAI has not released, with some results needing more than one attempt. The catalogue includes claimed progress toward the Riemann hypothesis and improvements to several well-known computer algorithms. That is a strikingly different story from the one OpenAI told a month earlier -- and it arrived at a company that, by its own admission, had just promised a group of outside mathematicians it would do this differently.
On September 8, OpenAI said a different internal model had produced a Lean-verified proof of a Navier-Stokes blowup result -- a documented sub-case of the Clay Mathematics Institute's Millennium Prize problem, not the full unforced version -- using roughly 10,000 coordinating agents over 88 hours and, by some estimates, millions of dollars of compute. That announcement was immediately overshadowed by NYU mathematician Tristan Buckmaster's published statement accusing OpenAI of learning about his and Anthropic researcher Levent Alpöge's related work first, then pressuring him to drop Alpöge as a co-author -- a dispute OpenAI's Sébastien Bubeck has denied and neither side's central technical claim has been independently reviewed. Three days later, 25 Fields Medalists signed an open letter titled "A Severe Misalignment of AI in Mathematics," warning that labs racing to announce solutions, without time for write-ups or attribution, risks breaking the field's chain of transmission. Timothy Gowers, one of the signatories, warned separately that mathematical literature could balloon in size while no human community actually understands what's in it. The same week, OpenAI withdrew its sponsorship of a math event at Caltech after researchers there publicly criticized the company.
OpenAI's two 2026 math claims, scoped
- 10,000 agents / 88 hrs · Sept. 8 claim
- One Navier-Stokes blowup result
Includes: A coordinated multi-agent swarm, per OpenAI's own account
Excludes: The model itself; independent confirmation of who originated the underlying technique - 1 agent / ~3 hrs avg · Oct. 6 release
- 372 result families (~720 files)
Includes: Lean-checked proof logic, per OpenAI; average compute time per problem
Excludes: The model, the exact prompts, and per-problem compute time -- the three things its own advisory group asked for
An advisory group OpenAI helped create, then didn't follow
The September fallout led directly to the body OpenAI now appears to have sidestepped. After OpenAI approached some mathematicians about forming an advisory board, nine of them -- including Fields Medalists Timothy Gowers and Martin Hairer, plus Ravi Vakil, Edward Witten, and Melanie Matchett Wood -- instead organized an independent Advisory Group on Mathematics and Artificial Intelligence, hosted at the Institute for Advanced Study in Princeton and announced around September 21-22. The IAS was explicit about the group's limits from the start: "we do not have decision making power at any AI company, and the responsibility for the decisions made by any company will rest with that company." OpenAI, for its part, said the group "will not be responsible for advising us on how to pace our internal progress on mathematics." Member Martin Hairer later described the IAS's own role as mostly logistical -- "so far mainly regarding IT, legal, and communications" -- underscoring how little formal authority the group was ever given, by design, over any lab's actual output.
The group's first assignment was narrower than pacing: advising OpenAI on how to coordinate the release of exactly the kind of mass result-dump that became the October 6 publication. Its stated recommendation -- reported consistently across coverage of the release -- was that OpenAI publish the model, the exact prompts used, and the compute time spent per problem. OpenAI released none of those three. It disclosed only aggregate figures: the ~3-hour average compute time per problem, and a claim that most results needed one prompt and one agent. Average is the word doing the work -- it tells a reader nothing about which of the 372 results took ten minutes and which took thirty hours, information a mathematician would need to judge how surprised to be by any single one, or whether a handful of easy wins are quietly padding a headline count.
Agents deployed, by claim
"A demonstration of power"
The Association for Human Mathematics -- a roughly 400-member group formed in August, separate from the IAS advisory body -- issued its own statement on October 7, posted as a guest entry on Terence Tao's blog. It is unambiguous. The release, it says, amounts to "a demonstration of power" rather than scholarship, pointing to "over 700 files at once" landing on a field with no mechanism to review them. It notes that OpenAI is currently defending lawsuits over plagiarism, copyright infringement, and trademark dilution, and argues the company disregarded the advisory group's own stated position that frontier labs shouldn't test advanced problems on internal models without a release plan already in place. Its recommendation: mathematicians should discontinue their work with OpenAI, and the field should return to a vision of science that "centers human understanding" over publication volume, with the public urged to view OpenAI's publication model "with due skepticism."
Tao's own position is more careful than the statement he hosted. He added no editorial comment of his own beyond a note that the post had been converted from another file format, signed simply "-- T." -- he has not endorsed a boycott. Writing separately, he described the release as marking the end of what he calls "Math 1.0" -- the era in which simply finding a solution to a named open problem was, by itself, the point of the work. He has also criticized the pace itself as "insane", warning that problems are being "solved" by a process with no interest in the surrounding field once the target is cleared, and that something is lost when the people doing the solving have no stake in training the next generation of mathematicians. That is a different complaint from AHM's: not that OpenAI lied, but that speed at this scale breaks the incentive structure that makes mathematics a *field* rather than a leaderboard.
I don't recommend anyone scramble to move their funds to new wallets today.
Not everyone in mathematics reads the release the same way. MIT's Andrew Sutherland has drawn the line at verification, not motive: single-agent, one-shot claims should be treated as unproven until an outside party can run the model and reproduce the result, and he wants "receipts" -- the prompts and the model -- before judging the work itself. That is a narrower complaint than AHM's: Sutherland isn't arguing OpenAI behaved badly, only that nobody should update their beliefs about the state of mathematics on the strength of a claim nobody else can test. University of Toronto's Daniel Litt takes the opposite tack on openness specifically: there is no good reason to keep correct answers secret once they exist, and publishing them, however abruptly, is a net gain for the field rather than an affront to it. Both object to different things; neither disputes that the underlying proof logic, where checked in Lean, appears sound. That split -- real researchers disagreeing about what the problem even is -- is itself evidence that "mathematicians are furious" oversimplifies a field that is, at minimum, divided into a verification camp and an openness camp, with AHM's boycott call sitting to one side of both, and Tao's "Math 1.0" framing sitting to a third side again.
Of the 372 results, independently replicated so far
A second alarm, from a field that wasn't even watching
The release traveled somewhere mathematicians weren't paying attention to. On October 7, theoretical computer scientist Scott Aaronson -- known for his work on quantum complexity, not a party to the math feud at all -- published a post titled "The Mathocalypse," noting that cryptography is conspicuously missing from OpenAI's 372 categories. You'd expect, he argued, that a model this capable at open problems would also take a swing at the number theory underneath RSA or the algebraic structure behind elliptic curve cryptography. Citing unnamed sources he trusts, Aaronson wrote that AI companies have begun, "gingerly and discreetly," testing whether their newest internal models can break "important cryptographic protocols and primitives" -- reasoning that if a model can do it, a lab would rather know first than let the rest of the world find out at the same time everyone else does.
That same day, Ethereum researcher Justin Drake posted that crypto holders should prepare for "bunker mode" -- citing OpenAI's math results directly as evidence that AI is advancing faster than expected on exactly the kind of problem that could threaten ECDSA, the elliptic-curve signature scheme securing most crypto wallets. His reasoning tracks Aaronson's: curves "carry rich structure, with room for fancy tricks like Schoof, Frobenius, pairings" -- unlike hash functions, which are deliberately engineered to minimize that kind of exploitable algebraic structure. This is a different threat model from the quantum-computing one Aaronson is better known for warning about via Shor's algorithm; Drake's and Aaronson's shared concern is that conventional mathematics, done faster by AI, could overturn the hardness assumptions curve-based cryptography relies on before any quantum computer is built at all. Drake's own recommendation was gradual, not panicked: "set in motion a controlled mass migration of assets to fresh addresses," sophisticated holders first, with an explicit warning that "a rushed migration would do more harm than good."
Vitalik Buterin responded the same week, agreeing the risk deserves attention while explicitly declining to endorse Drake's urgency. He went further than Drake in one respect: he said lattice-based cryptography -- the scheme most often proposed as the quantum-safe fallback, and itself a structured algebraic object not unlike the curves Drake named -- could take "serious hits" from AI-driven math over the next two years, which he said is a real reason Ethereum's long-term roadmap favors hash-based constructions over lattice- or curve-based ones. Hash functions have no comparable algebraic structure for a solver to exploit; that is close to the entire design philosophy behind choosing one over the other once you stop trusting that "nobody has found the trick yet" will hold indefinitely. But Buterin was explicit that no one should act out of fear. "I personally have lost more money in botched migrations than I have lost in all hacks combined," he said, framing a rushed, error-prone move to new wallets as the more probable near-term harm -- a user who sends funds to the wrong address while panic-migrating loses just as completely as one who gets hacked, and far more predictably. Haseeb Qureshi of Dragonfly called Drake's post "a very sober call" precisely because the threat it names is conventional mathematics overturning unproven hardness assumptions, not a quantum computer arriving on schedule -- a distinction that matters because the crypto industry has spent years preparing narratively for the quantum version of this story and comparatively none for the AI version, which nobody outside a small circle of researchers was war-gaming publicly before this week.
What's actually established
Strip away the framing on both sides and a short list of actually-confirmed facts remains, against a much longer list of things nobody outside the companies involved can yet check. That gap is the real story -- not whether OpenAI's mathematics is fake, and not whether a crypto wallet is unsafe tonight.
- OpenAI's model produced 372 results from one prompt to one agent each, nearly all on the first try.
- The 372 results are logically sound.
- OpenAI disregarded its own advisory group's guidance on how to release the results.
- AI could break the cryptography securing crypto wallets within the next year or two.
Neither side of that scorecard is static. OpenAI says it is working toward releasing the model; AHM's roughly 400 members could grow or shrink depending on whether any named university or mathematician actually withdraws from a collaboration rather than just signing a statement; and Drake's and Buterin's crypto timelines are, by their own framing, predictions about the next one to two years rather than settled fact. The strongest challenges to this piece's own framing come from exactly the two people already quoted pushing back hardest.
None of this requires OpenAI's claims to be false. Lean verification is a real check, and it is reportedly passing on most of the 372 results -- that is not nothing, and it is more than the Navier-Stokes claim had at the equivalent stage a month ago, when no formal verification tool had yet weighed in at all. What's missing in both threads is the same thing: a way for anyone outside OpenAI to independently confirm how a result was produced, at what cost, and under what conditions, before the rest of the world has to decide how much weight to put on it. The IAS advisory group asked for exactly that disclosure in September and didn't get it in October. Mathematicians without a seat on that board are left relying on OpenAI's own averages; a theoretical computer scientist and a group of crypto researchers reading the same release from entirely outside mathematics are left extrapolating a security timeline from a company that, on its own account, is not yet willing to show its work. Both groups are, in effect, asking the same question of the same release from opposite ends of the building.
(OpenAI has not announced a date for releasing the model behind the October results. The advisory group's own recommendations, and OpenAI's response to them, are not yet published in full -- what's public so far comes entirely from secondary reporting on both, which is why this piece treats each claim's level of confirmation separately rather than as a single verdict.) For now, the clearest fact in either story is the gap itself: a result set large enough to need an advisory group, an advisory group whose actual advice went unfollowed in the one instance it was asked to give it, and a second field entirely -- cryptography -- now watching the same release for a reason mathematicians never raised.
- OpenAI released 372 claimed math results (719-722 files) on Oct. 6, nearly all from one prompt to one unreleased model.
- This defied its own Institute for Advanced Study-hosted advisory group's call to publish the model, prompts, and compute time.
- The Association for Human Mathematics is urging mathematicians to stop working with OpenAI; Terence Tao is more measured.
- Scott Aaronson and Ethereum researcher Justin Drake separately warned AI could eventually threaten wallet cryptography.
- Caveat: no result has been independently replicated yet, and no practical cryptographic attack has been demonstrated.