On 6 October 2026 OpenAI published 722 math manuscripts in 372 result families, all produced by an internal model the public cannot use. The claims include the quasi-Riemann hypothesis, the Unique Games Conjecture and the Hodge conjecture for a whole class of shapes. AI solved math problems all year, but this was the largest single release, and the repository's own catalogue labels its computer checks "unchecked".

That gap between claimed and checked is the whole story. This guide lists every major problem AI solved in 2026, from Google DeepMind's Erdős study in January to OpenAI's release this week. Each result is placed on a four-step scale of how far its proof has really been verified. It also covers the two questions everyone is asking: did AI solve Navier-Stokes, and did it prove the Riemann hypothesis?

The Key Takeaways

  • The release: OpenAI posted 722 manuscripts in 372 families from about 4,000 problems it posed to an unreleased model, at roughly three hours of ChatGPT Pro compute per result.
  • The checking: OpenAI links Lean proofs for 235 of the 372 families, and its formal catalogue gives its own review status as "unchecked".
  • Riemann: not proved. OpenAI claims a weaker statement, that no zeros lie where the real part exceeds 7/8. Nobody outside OpenAI has confirmed it.
  • Navier-Stokes: the Clay Institute says the problem "has apparently been settled", but only in the version that allows an external force.
  • Best-checked AI results: the unit distance disproof (May), the cycle double cover proof (July) and the Jacobian counterexample (July). Outside mathematicians checked all three.

What OpenAI Released on 6 October 2026

Vom Herausgeber

Jedes KI-Modell in einer App

Fello AI vereint GPT-6, Claude 5.5, Gemini 3.8, Grok 4.7 und mehr in einer nativen App für Mac und iPhone.

Jetzt herunterladen!

OpenAI released a public GitHub repository, openai/math, holding 722 manuscripts grouped into 372 families of related results. According to the repository's README, the model was "posed approximately 4,000 problems", and each family is a principal result plus companion papers, consequences or alternative proofs. The New York Times counted the release as 377 results. The 372 figure comes from OpenAI's own catalogue.

OpenAI announced the release from its main account, naming the advisory group it consulted:

The Headline Claims

OpenAI's overview document describes each family in one paragraph. The biggest names in the 6 October release are below, with whether OpenAI links a Lean proof for each one.

Claimed resultWhat it saysLean entry?
Quasi-Riemann hypothesisNo Dirichlet L-function, including the Riemann zeta function, has zeros where the real part of s exceeds 7/8Yes
Unique Games ConjectureSubhash Khot's conjecture on the limits of approximating hard optimization problems, provedYes
Mahler conjecturesSymmetric and nonsymmetric versions, in every dimensionYes
Kaplansky zero-divisor conjectureDisproved by a constructed counterexample groupYes
Rational Hodge conjecture for CM abelian varietiesA special case of a Millennium Prize Problem, for one family of geometric objectsNo
Kakeya conjecturesThe maximal conjecture in three dimensions and the dimension conjecture in fourNo
Free group factorsAll nonabelian free group factors are isomorphicYes
Birch and Swinnerton-Dyer formulaThe full formula for elliptic curves with Selmer corank zero or oneNo

OpenAI says the "vast majority" of results came from one fixed procedure on the unreleased model. An OpenAI spokesperson told Scientific American that almost every result came from a single prompt to a single agent, though some may have taken multiple attempts. Two results were exceptions to the procedure: the zeta zero-free region and the Hodge result. The write-up of the 11/12 zero-free region was also edited by humans for readability.

How Much of the Release a Computer Has Checked

OpenAI's manuscript map links Lean formalizations for 235 of the 372 families, backed by 405 machine-checkable challenge files. The separate catalogue file lists 162 papers, sets its scope as "Partial progress." and gives its review status as "unchecked". The README warns that "some of the unformalized results could have issues".

A Lean entry can also cover less than the paper claims. The documentation for the free group factor result says one of the paper's conclusions is a further consequence "rather than a separate selected statement" in the formal proof.

The coverage is uneven across the headline claims, as the table shows. For the Hodge, Kakeya and Birch and Swinnerton-Dyer results, the only evidence today is a PDF written by the model.

The Fello AI Proof Ladder: How to Tell a Solved Problem From a Claim

The Fello AI proof ladder is a four-step scale for how far an AI math result has been verified. It runs from a lab's announcement to acceptance by a journal or prize body. Each step is harder to reach than the one below it, and each one rules out a different kind of error.

RungWhat happenedWhat it rules outWhat it does not rule out
1. ClaimedThe lab published a proofNothingAny error
2. Machine-checkedA Lean proof compilesLogical gaps in the formal proofA formal statement that differs from the real problem; lack of novelty or importance
3. Checked by mathematiciansNamed outside experts read and confirmed itMisreading of the problem; most hidden errorsDisputes over credit; unpublished subtleties
4. AcceptedA journal published it or a prize body ruledAlmost everythingVery little

Rung 2 is where most confusion starts. Lean is a programming language in which a proof only compiles when every step follows from the previous one, so a compiled proof is logically sound. But Lean checks the statement as someone typed it, not the problem as mathematicians mean it. In September, Tom Adamczewski of Epoch AI posted GPT-6 Astra's Lean disproof of the 1930 Köthe conjecture. He wrote that it held "at least as stated in the Formal Conjectures repo". He added that he was "not competent to judge" whether the problem had been misformalized.

Rung 4 is slow by design. The Clay Mathematics Institute calls its own Millennium Prize process "deliberately unhurried".

AI Solved Math Problems in 2026: The Full Record

AI solved math problems in 2026 at three labs, and most of the best-checked results are older than this week's headlines. The table lists the major results in date order, with the highest rung each one has reached on the Fello AI proof ladder as of 7 October 2026.

DateProblemLab and modelRungEvidence
Jan 2026700 "open" Erdős problems surveyedGoogle DeepMind, Gemini313 addressed: 5 seemingly new solutions, 8 already solved in the literature. Human experts graded each one.
May 2026Erdős unit distance conjecture (1946), disprovedOpenAI, internal model3Human-verified write-up by nine mathematicians including Timothy Gowers and Melanie Wood
Jul 2026Cycle double cover conjecture, provedOpenAI3Independent expositions by Jim Geelen and Sang-il Oum
Jul 2026Jacobian conjecture, false in three dimensionsLevent Alpöge with Claude Fable 53Short enough to check by hand or in ten lines of code
Aug 2026Ten results, including the first non-sofic group and three Erdős problemsOpenAI, internal version of GPT-6 Astra2Lean certificates for all ten, published on GitHub
Aug 2026Share of zeta zeros on the critical line raised from 41.6% to 67.2%Anthropic, unreleased Claude3Lean proof passes the comparator tool; experts Brian Conrey and Dan Goldston examined the paper
Sep 2026Köthe conjecture (1930), disproved in LeanEpoch AI with GPT-6 Astra2Lean proof; poster flagged possible misformalization
Sep 2026Navier-Stokes Millennium Problem, forced versionOpenAI, internal model3Lean proof; Clay says it "has apparently been settled", prize ruling pending
Oct 2026Quasi-Riemann hypothesisOpenAI, internal model2Lean entry in a catalogue marked "unchecked"
Oct 2026Unique Games ConjectureOpenAI, internal model2Lean entry in the same catalogue
Oct 2026Hodge conjecture for CM abelian varietiesOpenAI, internal model1Model-written paper only
Oct 2026Kakeya conjectures in 3D and 4DOpenAI, internal model1Model-written paper only

What the Rung 3 Results Have in Common

Two of the mathematicians who checked the unit distance disproof rank it above every earlier AI result. In May, Daniel Litt called it "the unique interesting result produced autonomously by AI so far". Timothy Gowers wrote in commentary solicited by OpenAI that "No previous AI-generated proof has come close" to the standard of a top journal. OpenAI itself is modest about the method. Sébastien Bubeck, who leads OpenAI's math work, told Scientific American: "The model did not invent something fundamentally new that nobody saw coming."

The DeepMind study points the same way. Its authors concluded that the 'Open' status of the Erdős problems they resolved came "through obscurity rather than difficulty". They also warned of "subconscious plagiarism" by AI, meaning a model reproducing a published idea without knowing it. Anthropic said the same of its own zeta result: "We don't expect that the techniques Claude used will lead to proving the Riemann hypothesis."

Did AI Solve Navier-Stokes?

Apparently yes for the version Clay wrote in 2000, and no for the version most fluid experts care about. OpenAI's internal model proved that a smooth fluid starting at rest can develop a singularity in finite time when a smooth external force pushes on it. That is statement C in the official Millennium Prize problem.

The Clay Mathematics Institute's 11 September statement says the problem "has apparently been settled", and promises updates on the prize. OpenAI says it does not intend to claim the $1 million.

The dispute is about the force.

Most experts study the problem with no external force, so that any blowup comes from the fluid itself. "The Clay problem is settled, but the main problem for the Navier-Stokes equations is not," University of Chicago mathematician Luis Silvestre told Scientific American. On 17 September, Peter Constantin, Mihaela Ignatova and Vlad Vicol posted a paper showing that OpenAI's construction cannot work if the force vanishes near the singular point.

The result also arrived amid a credit dispute. NYU mathematician Tristan Buckmaster and Anthropic's Levent Alpöge had published a related Euler blowup first. OpenAI's own post recognizes "the priority of their work on forced Euler" and says its model never saw their work.

Did AI Prove the Riemann Hypothesis?

No. The Riemann hypothesis says every non-trivial zero of the zeta function has real part exactly 1/2. OpenAI's claimed result is that no zeros have real part greater than 7/8. That rules out a strip of the plane, but it says nothing about the region between 1/2 and 7/8 where the real question lives.

Two AI results in 2026 sit close to the Riemann hypothesis, and neither proves it. In August, an unreleased research version of Anthropic's Claude raised the proven share of zeros on the critical line from 41.6% to 67.2%. It had been asked to "take a real stab" at the full hypothesis and found the bound along the way.

In October, OpenAI claimed the quasi-Riemann hypothesis at 7/8, with a companion proof at 11/12. It has a Lean entry but no outside review yet.

Scientific American listed the 7/8 claim among OpenAI's results as "actual progress toward" the Riemann hypothesis. If it survives review, any zero off the critical line would have to sit between 1/2 and 7/8. But any headline claiming AI proved the Riemann hypothesis is wrong about both results.

Check an AI Math Result Yourself in Ten Lines of Python

The Jacobian conjecture counterexample is the easiest AI result of 2026 for a reader to verify on a laptop. The conjecture, stated in its general form in 1939, says a polynomial map whose Jacobian determinant is a non-zero constant can always be reversed. Levent Alpöge, a mathematician at Anthropic, posted a three-dimensional map that breaks it, found with Claude Fable 5:

Run the Check

A reversible map can never send two different points to the same place. This one sends three. Install SymPy with pip install sympy, then run:

from sympy import symbols, Matrix, expand, Rational
x, y, z = symbols('x y z')
F = Matrix([
    (1 + x*y)**3 * z + y**2 * (1 + x*y) * (4 + 3*x*y),
    y + 3*x*(1 + x*y)**2 * z + 3*x*y**2 * (4 + 3*x*y),
    2*x-3*x**2*y-x**3*z,
])
print("Jacobian determinant:", expand(F.jacobian([x, y, z]).det()))
for p in [(0, 0, Rational(-1, 4)), (1, Rational(-3, 2), Rational(13, 2)), (-1, Rational(3, 2), Rational(13, 2))]:
    print(p, "->", tuple(F.subs({x: p[0], y: p[1], z: p[2]})))

We ran it on 7 October 2026. The determinant printed as -2, and all three points printed as (-1/4, 0, 0). That is the whole proof: a constant non-zero determinant and three points merged into one. The hard part was finding the formula, not checking it, which is why this result sits on rung 3 while far grander claims sit on rung 1.

Why Mathematicians Object to How AI Solved Math Problems

The loudest objections to how AI solved math problems this year are about release, not correctness. The Advisory Group on Mathematics and Artificial Intelligence (AGMAI) is nine mathematicians hosted at the Institute for Advanced Study. It published release guidelines on 29 September based on more than 600 survey replies. Its members include Timothy Gowers, Martin Hairer, Edward Witten and Melanie Wood.

What the Advisory Group Asked For, and What OpenAI Did

OpenAI's 6 October release meets some of the group's requests and skips others. The comparison below uses AGMAI's text, OpenAI's README and blog post, and reporting by Scientific American and the New York Times.

AGMAI recommendationOpenAI's 6 October release
Stop testing advanced problems on proprietary modelsNot followed. The results come from an unreleased internal model
Publish the model name, prompts, reasoning, time and cost for each resultPartly. 10 reasoning summaries and an average of about three hours of Pro compute; no prompts
Deposit results in a repository no AI lab controlsNot yet. GitHub repo under OpenAI's account; OpenAI says it is exploring community-hosted options
Formalize proofs, or state the formalization status clearlyPartly. Lean links for 235 of 372 families; none for the Hodge, Kakeya or Birch and Swinnerton-Dyer results
Report how many problems were tried and failed, and how they were chosenPartly. About 4,000 problems posed
Fund work on human understanding of the resultsPromised. OpenAI says it will fund workshops and conferences

An OpenAI spokesperson told Scientific American the company takes the guidelines seriously, but added that OpenAI is not bound by them. AGMAI's own statement on 6 October said its role "should not be interpreted as a judgment of the impact of these results or an endorsement". It called the release "the beginning, not the completion, of the process of human understanding".

The Fields Medallists and Terence Tao

The sharpest criticism came from the top of the field. On 11 September, Fields Medallists published a declaration that "the goals of the AI companies and the goals of the mathematical community are severely misaligned". It warned that mass-produced "true/false" statements "could destroy fertile ground instead of breathing life into new ideas". Scientific American counted 25 signatories at launch; the page now lists 28, including Terence Tao, Peter Scholze and Maryna Viazovska.

Tao appeared in an OpenAI promotional video earlier in the year. On 9 September he wrote on his blog that the community should "reject irresponsible and unsustainable usages of AI technology that only serve to advance nominal goals rather than the true underlying goals of the field".

Others welcome the flood. "If we want to know the answers to these math questions, I see no reason why we should ask the company to keep them secret from us," University of Toronto mathematician Daniel Litt told Scientific American. MIT's Andrew Sutherland put the skeptical case in five words: "We should ask for receipts."

The disagreement echoes a wider argument about AI labs racing ahead of the people who must check their work. Our guide to AI misalignment covers the safety side of that argument.

Can You Use the Model That Solved These Problems?

No. The model behind the Navier-Stokes proof and the 6 October release is an internal OpenAI system that has not been released. OpenAI says it is "working to responsibly release the model" but has given no date. It described the Navier-Stokes model as "significantly more capable than GPT‑6 Astra", OpenAI's flagship public model, covered in our GPT-6 Astra review.

Public models are still capable of research-grade work. Alpöge found the Jacobian counterexample with Claude Fable 5, which had reached the public only weeks earlier, according to ScienceDaily. Our Claude Fable 5 overview covers what that model can do. Epoch AI's Köthe disproof used GPT-6 Astra. For everyday math, from homework to checking a proof sketch, our best AI for math guide compares the public models task by task.

Compare the Public Models Side by Side

The models disagree with each other on hard math, and the disagreement is useful. Asking two of them the same question is the fastest way to spot a step one of them made up. Fello AI puts GPT, Claude, Gemini, Grok and DeepSeek in one app on Mac, iPhone and iPad. You can put the same question to each model without paying for a separate subscription to each. It has a free tier, and the full plan costs $9.99 a month.

What Happens Next

The verdict on the 6 October release will come from mathematicians, not from OpenAI, and it will take months. For now, the solid AI results of 2026 are the ones named outside experts have read: the unit distance disproof, the cycle double cover proof, the Jacobian counterexample, Claude's zeta bound and Navier-Stokes as Clay wrote it. The quasi-Riemann hypothesis and the Unique Games proof have a computer check behind them and nothing else yet. The Hodge and Kakeya claims are papers.

The best model for releasing AI math already exists, and OpenAI built it. In May, the unit distance disproof arrived with commentary from outside mathematicians on day one, and nine of them then published a human-verified version. When the next headline says AI solved a famous problem, ask which rung it is on. For how fast these systems are improving overall, our AGI prediction tracker scores every forecast. Our AI benchmarks guide explains the math benchmarks that OpenAI says its models saturated before it moved on to open problems.

FAQ

Has AI solved a Millennium Prize Problem?

Possibly one, in a narrow sense. The Clay Mathematics Institute says OpenAI's Navier-Stokes result means the problem "has apparently been settled" in its official form, which allows an external force. Fluid mathematicians such as Luis Silvestre say the unforced version, the one most experts study, is still open. Clay has not ruled on the prize.

Did AI prove the Riemann hypothesis?

No. OpenAI claims the quasi-Riemann hypothesis, that zeta has no zeros with real part above 7/8, which is a much weaker statement and is not yet reviewed. Anthropic's Claude raised the share of zeros proven to lie on the critical line to 67.2%, which also falls short of the full hypothesis.

How many math problems has AI solved?

Nobody has a verified count. OpenAI alone claims 372 result families from its 6 October release, but most are unreviewed. Far fewer have been checked by named outside mathematicians. Google DeepMind's study alone had experts grade 13 Erdős problems, and a handful of famous conjectures, such as unit distance and cycle double cover, have published human write-ups.

What is Lean, and does a Lean proof mean a result is true?

Lean is a programming language in which a proof only compiles if every step is logically valid. A compiled Lean proof means the formal statement is true. It does not guarantee that the formal statement matches the problem mathematicians meant, or that the result is new or important.

Can I use the AI model that solved these problems?

Not yet. OpenAI's math results came from an unreleased internal model, and the company has given no release date. Public models such as GPT-6 Astra and Claude Fable 5 have produced checked results of their own, including the Köthe disproof and the Jacobian counterexample.