Google DeepMind’s AlphaProof Nexus paper published in Science
LLM + Lean agent resolved 9 of 353 open Erdős problems and 44 of 492 OEIS conjectures
Result confirmed
Importance: major (4 of 5)The takeaway
On 8 Oct 2026 Science published Google DeepMind’s AlphaProof Nexus paper (“Advancing mathematics research with AI-driven formal proof search”, doi 10.1126/science.aej2213), first posted on arXiv on 21 May 2026.
Status
- Claim
Result confirmed
- Our reporting
- High confidence
- Verification
- Peer-reviewed
- Importance
- Major (4 of 5)
- Last verified
- 10 October 2026
Your AI and this story
- GPT-6 Astra161 days after its cutoff
- Claude Opus 5.5100 days after its cutoff
- Gemini 3.8 Flash191 days after its cutoff
- Grok 4.7130 days after its cutoff
None of these four assistants can know about it. The closest, Claude Opus 5.5, stops 100 days before it.
Key facts
- Science, 8 Oct 2026, doi 10.1126/science.aej2213 (received 27 May, accepted 17 Aug 2026, per press reports); preprint arXiv 2605.22763 (v1 21 May 2026); 21 authors led by George Tsoukalas, Anton Kovsharov, Sergey Shirobokov; correspondence Pushmeet Kohli and Swarat Chaudhuri
- Erdős problems: 9 of 353 attempted (~2.5%), counted by formal sub-question: #12(i), #12(ii), #125 (variant), #138 (variant), #152, #26 (Tenenbaum variant), #741(i), #741(ii), #846 (Lean files in github.com/google-deepmind/alphaproof-nexus-results)
- The ‘two questions open for 56 years’ are Erdős #12(i) and #12(ii), posed by Erdős and Sárközy in 1970 (paper, Sec. 3)
- OEIS: Gemini autoformalized 492 open OEIS conjectures; the agent proved 44 that manual review found correctly formalized and previously unproven. Test lemmas checking the first terms of each sequence guarded against misformalization
- Comparison: the later OEIS Open benchmark (arXiv 2608.11941) reports 147 of 492 formalised open OEIS conjectures resolved by newer models at $50/attempt; OpenAI’s 6 Oct 2026 release lists 722 manuscripts, mostly not formally verified (see related entries)
- Deployment: proved an exact O(1/t) last-iterate rate for Anchored Gradient Descent-Ascent in min-max convex-concave optimization while also discovering a new step-size schedule; solved 2 of 4 open Hilbert-function problems with Gergely Bérczi, incl. Zanello’s log-concavity conjecture for pure O-sequences of codimension 3 and type 2; helped on a Ben Green list problem; quantum optics problems with Mario Krenn
- Agents: a ‘basic’ agent (prover subagents + Lean compiler feedback) and a ‘full-featured’ agent with evolutionary coordination, Elo-rated sketches (Gemini 3.0 Flash raters) and AlphaProof as a tool; the basic agent solved all 9 Erdős problems but cost more on the hardest
- Science vol. 394, issue 6820, pp. 234-239 (Research Article). Same issue: Perspective ‘AI for research mathematics has arrived’ by Jeremy Avigad (CMU) and Matthew Ballard (U. South Carolina), pp. 166-167 (per EurekAlert: even unsuccessful AI proof attempts can help mathematicians understand problems), and an ‘Expert Voices’ piece by Emily Riehl, ‘AI has solved many math problems, but it has not solved math’
Show 2 more
- erdosproblems.com credits DeepMind on #12, #26, #125, #138, #152, #741, #846; #741 and #846 were found independently by an internal OpenAI model too; for #846 the page notes a counterexample also follows from earlier work of Reiher, Rödl et al.
- Viral 2nd wave: @Dr_Singularity post on 10 Oct 2026 (~45.6k views at fetch time) presented the result as new
What happened
On 8 October 2026 Science published Google DeepMind’s paper on AlphaProof Nexus, a framework of LLM agents that search for formal proofs in Lean. The preprint (arXiv 2605.22763) had appeared on 21 May 2026 and was covered then by the tech press. Publication in Science, after peer review, brought a second wave: an EurekAlert release, expert reactions from the Spanish Science Media Centre, a Science Perspective by Jeremy Avigad and Matthew Ballard, an “Expert Voices” piece by Emily Riehl in the same issue, a Pushmeet Kohli thread on X, and on 10 October a viral summary by @Dr_Singularity.
The system pairs Gemini 3.1 Pro prover subagents with the Lean compiler. Every proposed proof is machine-checked, and failed attempts feed back into the next try. The paper compares a basic agent (independent subagents plus compiler feedback) with a full-featured agent, which coordinates subagents with an evolutionary algorithm, ranks proof sketches with Gemini 3.0 Flash raters and can call AlphaProof (the 2024 RL prover) as a tool.
Results (from the paper):
- Erdős problems: 9 of 353 attempted, autonomously, at a few hundred dollars of inference per problem. The 9 are formal sub-questions: #12(i), #12(ii), #125, #138, #152, #26, #741(i), #741(ii) and #846, several in variant forms. The two “open for 56 years” are #12(i) and (ii), posed by Erdős and Sárközy in 1970. In a post-hoc ablation, the basic agent also solved all 9, at higher cost on the hardest ones.
- OEIS: Gemini autoformalized 492 open conjectures from the Online Encyclopedia of Integer Sequences. The agent proved 44 that manual review judged correctly formalized and previously unproven.
- Research deployments: an exact O(1/t) last-iterate rate for Anchored Gradient Descent-Ascent in min-max convex-concave optimization, found together with a new step-size schedule. With Gergely Bérczi (Aarhus), it solved 2 of 4 open problems on Hilbert functions, including Zanello’s log-concavity conjecture for pure O-sequences in the (codimension 3, type 2) case. It also helped resolve a problem from Ben Green’s list, found misformalizations in the literature and supported quantum optics work with Mario Krenn.
Why it matters
This is a peer-reviewed, fully formal account of LLM agents solving open research problems at scale and at low cost. Its main finding, that a simple generate-and-check loop does about as well as a complex agent, supports the view that model capability, not agent engineering, now drives AI mathematical discovery.
Verification and reactions
- Lean-verified and peer-reviewed. All proofs are public Lean files (github.com/google-deepmind/alphaproof-nexus-results), and erdosproblems.com credits DeepMind on the pages of the solved problems. Science’s article pages were not readable for this entry (HTTP 403). Volume and pages come from Science’s RSS table of contents; the received and accepted dates come from press summaries.
- Not all fully new. On erdosproblems.com, #741 and #846 are credited to DeepMind and an internal OpenAI model, found independently. The #846 page notes that a counterexample also follows from a theorem of Reiher, Rödl et al. Several solutions settle variants or single parts of a problem.
- Critics. Anatol Wegner’s blog (31 May, written with Gemini assistance; not peer-reviewed) argues that some “open” problems had prior literature and that the agent exploited loose formalizations. These are his claims and were not checked here. The Perspective and Riehl’s piece were seen only as titles in Science’s table of contents (the pages return HTTP 403 to us); their text is not summarised here. In the SMC Spain reactions, Josep Curto (UOC) called the system “a severe case of over-engineering”, since the basic agent matched it. Javier Aramayona (ICMAT-CSIC) warned of results “known to be true” that “no one fully understands”. Pablo Haya Coll (UAM) and Teodoro Calonge (Valladolid) were positive.
- In May, WION reported Demis Hassabis saying AGI is still far away despite the result, and a video posted to Hacker News was titled “Hassabis said solving Erdős isn’t real invention” (not checked against the video).
Comparison
Measured by count, AlphaProof Nexus is now well behind later systems. The OEIS Open benchmark (arXiv 2608.11941) reports 147 of the same 492 formalised OEIS conjectures resolved at $50 per attempt. OpenAI’s 6 Oct release lists 722 manuscripts, but only 162 are Lean-formalized and their review is marked “unchecked”. The 44 OEIS and 9 Erdős results were obtained with a May 2026 model. What sets this paper apart is that every result is formally verified and the work passed peer review in a top general-science journal.
Sources
16 sources from 11 sites. Numbers match the chips in the text.
16 sources: 4 primary, 5 press, 7 reactions
Primary
- Science: Advancing mathematics research with AI-driven formal proof search (doi 10.1126/science.aej2213)science.org, paper
- Science Perspective: AI for research mathematics has arrived (Avigad & Ballard)science.org, paper
- arXiv 2605.22763: Advancing Mathematics Research with AI-Driven Formal Proof Searcharxiv.org, paper
- GitHub: google-deepmind/alphaproof-nexus-results (Lean proofs + prose proofs)github.com, code
Press
- Science Expert Voices: AI has solved many math problems, but it has not solved math (Emily Riehl)science.org, press
- EurekAlert (AAAS): Introducing AlphaProof Nexus (8 Oct 2026)eurekalert.org, press
- Science Media Centre Spain: expert reactions (8 Oct 2026)sciencemediacentre.es, press
- The Decoder: AlphaProof Nexus solves decades-old math problems for a few hundred dollars (May 2026)the-decoder.com, press
- WION: Google AI solves decades-old maths problems but DeepMind CEO says AGI is still far away (May 2026)embed.wionews.com, press
Reactions
- Pushmeet Kohli on X: thread on DeepMind and mathematics, paper in Science today (8 Oct 2026)x.com, discussion
- Pushmeet Kohli on X: ‘Our technical paper on AlphaProof Nexus appears in Science today’x.com, discussion
- Dr Singularity on X: viral summary (10 Oct 2026)x.com, discussion
- erdosproblems.com #12 (DeepMind construction)erdosproblems.com, discussion
- erdosproblems.com #846 (DeepMind and OpenAI, independently)erdosproblems.com, discussion
- Anatol Wegner: critical analysis of the AlphaProof Nexus paper (31 May 2026)buttondown.com, discussion
- Hacker News discussion (May 2026)news.ycombinator.com, discussion
Changes
- Filed (Science publication 8 Oct; the May arXiv preprint had only been mentioned in passing in OpenAI model disproves Erdős’s 80-year-old unit distance conjecture). The viral X post of 10 Oct was the trigger.