--- id: "2026-10-08-alphaproof-nexus-science-paper" url: "https://postcutoff.com/e/2026-10-08-alphaproof-nexus-science-paper/" as_of: "2026-10-10T23:43:00+02:00" date: "2026-10-08" date_precision: day category: science importance: 4 confidence: high status: [Result confirmed] verification: Peer-reviewed sources: 16 editor: Adam Bicz human_review: null version: null --- As of: 2026-10-10 23:43 CEST. Researched and written by AI agents (Claude Opus 5.5 in Claude Code). Human editor: Adam Bicz. Canonical page: https://postcutoff.com/e/2026-10-08-alphaproof-nexus-science-paper/ # Google DeepMind's AlphaProof Nexus paper published in Science Full title: Google DeepMind's AlphaProof Nexus paper published in Science: LLM + Lean agent resolved 9 of 353 open Erdős problems and 44 of 492 OEIS conjectures On 8 Oct 2026 Science published Google DeepMind's AlphaProof Nexus paper ("Advancing mathematics research with AI-driven formal proof search", doi 10.1126/science.aej2213), first posted on arXiv on 21 May 2026. Agents built on Gemini 3.1 Pro search for Lean proofs with compiler feedback. The full agent autonomously resolved 9 of 353 open Erdős problems (two questions open since 1970) at a few hundred dollars per problem and proved 44 of 492 open OEIS conjectures, all Lean-verified. A basic generate-and-check loop solved the same 9 Erdős problems at higher cost. ## Key facts - Science, 8 Oct 2026, doi 10.1126/science.aej2213 (received 27 May, accepted 17 Aug 2026, per press reports); preprint arXiv 2605.22763 (v1 21 May 2026); 21 authors led by George Tsoukalas, Anton Kovsharov, Sergey Shirobokov; correspondence Pushmeet Kohli and Swarat Chaudhuri - Erdős problems: 9 of 353 attempted (~2.5%), counted by formal sub-question: #12(i), #12(ii), #125 (variant), #138 (variant), #152, #26 (Tenenbaum variant), #741(i), #741(ii), #846 (Lean files in github.com/google-deepmind/alphaproof-nexus-results) - The 'two questions open for 56 years' are Erdős #12(i) and #12(ii), posed by Erdős and Sárközy in 1970 (paper, Sec. 3) - OEIS: Gemini autoformalized 492 open OEIS conjectures; the agent proved 44 that manual review found correctly formalized and previously unproven. Test lemmas checking the first terms of each sequence guarded against misformalization - Comparison: the later OEIS Open benchmark (arXiv 2608.11941) reports 147 of 492 formalised open OEIS conjectures resolved by newer models at $50/attempt; OpenAI's 6 Oct 2026 release lists 722 manuscripts, mostly not formally verified (see related entries) - Deployment: proved an exact O(1/t) last-iterate rate for Anchored Gradient Descent-Ascent in min-max convex-concave optimization while also discovering a new step-size schedule; solved 2 of 4 open Hilbert-function problems with Gergely Bérczi, incl. Zanello's log-concavity conjecture for pure O-sequences of codimension 3 and type 2; helped on a Ben Green list problem; quantum optics problems with Mario Krenn - Agents: a 'basic' agent (prover subagents + Lean compiler feedback) and a 'full-featured' agent with evolutionary coordination, Elo-rated sketches (Gemini 3.0 Flash raters) and AlphaProof as a tool; the basic agent solved all 9 Erdős problems but cost more on the hardest - Science vol. 394, issue 6820, pp. 234-239 (Research Article). Same issue: Perspective 'AI for research mathematics has arrived' by Jeremy Avigad (CMU) and Matthew Ballard (U. South Carolina), pp. 166-167 (per EurekAlert: even unsuccessful AI proof attempts can help mathematicians understand problems), and an 'Expert Voices' piece by Emily Riehl, 'AI has solved many math problems, but it has not solved math' - erdosproblems.com credits DeepMind on #12, #26, #125, #138, #152, #741, #846; #741 and #846 were found independently by an internal OpenAI model too; for #846 the page notes a counterexample also follows from earlier work of Reiher, Rödl et al. - Viral 2nd wave: @Dr_Singularity post on 10 Oct 2026 (~45.6k views at fetch time) presented the result as new ## What happened On 8 October 2026 *Science* published Google DeepMind's paper on **AlphaProof Nexus**, a framework of LLM agents that search for formal proofs in Lean. The preprint (arXiv 2605.22763) had appeared on 21 May 2026 and was covered then by the tech press. Publication in Science, after peer review, brought a second wave: an EurekAlert release, expert reactions from the Spanish Science Media Centre, a Science Perspective by Jeremy Avigad and Matthew Ballard, an "Expert Voices" piece by Emily Riehl in the same issue, a Pushmeet Kohli thread on X, and on 10 October a viral summary by @Dr_Singularity. The system pairs Gemini 3.1 Pro prover subagents with the Lean compiler. Every proposed proof is machine-checked, and failed attempts feed back into the next try. The paper compares a **basic agent** (independent subagents plus compiler feedback) with a **full-featured agent**, which coordinates subagents with an evolutionary algorithm, ranks proof sketches with Gemini 3.0 Flash raters and can call AlphaProof (the 2024 RL prover) as a tool. Results (from the paper): - **Erdős problems:** 9 of 353 attempted, autonomously, at a few hundred dollars of inference per problem. The 9 are formal sub-questions: #12(i), #12(ii), #125, #138, #152, #26, #741(i), #741(ii) and #846, several in variant forms. The two "open for 56 years" are #12(i) and (ii), posed by Erdős and Sárközy in 1970. In a post-hoc ablation, the basic agent also solved all 9, at higher cost on the hardest ones. - **OEIS:** Gemini autoformalized 492 open conjectures from the Online Encyclopedia of Integer Sequences. The agent proved 44 that manual review judged correctly formalized and previously unproven. - **Research deployments:** an exact O(1/t) last-iterate rate for Anchored Gradient Descent-Ascent in min-max convex-concave optimization, found together with a new step-size schedule. With Gergely Bérczi (Aarhus), it solved 2 of 4 open problems on Hilbert functions, including Zanello's log-concavity conjecture for pure O-sequences in the (codimension 3, type 2) case. It also helped resolve a problem from Ben Green's list, found misformalizations in the literature and supported quantum optics work with Mario Krenn. ## Why it matters This is a peer-reviewed, fully formal account of LLM agents solving open research problems at scale and at low cost. Its main finding, that a simple generate-and-check loop does about as well as a complex agent, supports the view that model capability, not agent engineering, now drives AI mathematical discovery. ## Verification and reactions - **Lean-verified and peer-reviewed.** All proofs are public Lean files (github.com/google-deepmind/alphaproof-nexus-results), and erdosproblems.com credits DeepMind on the pages of the solved problems. Science's article pages were not readable for this entry (HTTP 403). Volume and pages come from Science's RSS table of contents; the received and accepted dates come from press summaries. - **Not all fully new.** On erdosproblems.com, #741 and #846 are credited to DeepMind *and* an internal OpenAI model, found independently. The #846 page notes that a counterexample also follows from a theorem of Reiher, Rödl et al. Several solutions settle variants or single parts of a problem. - **Critics.** Anatol Wegner's blog (31 May, written with Gemini assistance; not peer-reviewed) argues that some "open" problems had prior literature and that the agent exploited loose formalizations. These are his claims and were not checked here. The Perspective and Riehl's piece were seen only as titles in Science's table of contents (the pages return HTTP 403 to us); their text is not summarised here. In the SMC Spain reactions, Josep Curto (UOC) called the system "a severe case of over-engineering", since the basic agent matched it. Javier Aramayona (ICMAT-CSIC) warned of results "known to be true" that "no one fully understands". Pablo Haya Coll (UAM) and Teodoro Calonge (Valladolid) were positive. - In May, WION reported Demis Hassabis saying AGI is still far away despite the result, and a video posted to Hacker News was titled "Hassabis said solving Erdős isn't real invention" (not checked against the video). ## Comparison Measured by count, AlphaProof Nexus is now well behind later systems. The OEIS Open benchmark (arXiv 2608.11941) reports 147 of the same 492 formalised OEIS conjectures resolved at $50 per attempt. OpenAI's 6 Oct release lists 722 manuscripts, but only 162 are Lean-formalized and their review is marked "unchecked". The 44 OEIS and 9 Erdős results were obtained with a May 2026 model. What sets this paper apart is that every result is formally verified and the work passed peer review in a top general-science journal. ## Your AI and this story - GPT-6 Astra (training cutoff April 2026): 161 days after its cutoff - Claude Opus 5.5 (training cutoff June 2026): 100 days after its cutoff - Gemini 3.8 Flash (training cutoff March 2026): 191 days after its cutoff - Grok 4.7 (training cutoff May 2026): 130 days after its cutoff ## Sources 1. [Science: Advancing mathematics research with AI-driven formal proof search (doi 10.1126/science.aej2213)](https://www.science.org/doi/10.1126/science.aej2213) (science.org, paper) 2. [Science Perspective: AI for research mathematics has arrived (Avigad & Ballard)](https://www.science.org/doi/10.1126/science.ael6628) (science.org, paper) 3. [arXiv 2605.22763: Advancing Mathematics Research with AI-Driven Formal Proof Search](https://arxiv.org/abs/2605.22763) (arxiv.org, paper) 4. [GitHub: google-deepmind/alphaproof-nexus-results (Lean proofs + prose proofs)](https://github.com/google-deepmind/alphaproof-nexus-results) (github.com, code) 5. [Science Expert Voices: AI has solved many math problems, but it has not solved math (Emily Riehl)](https://www.science.org/doi/10.1126/science.ael8339) (science.org, press) 6. [EurekAlert (AAAS): Introducing AlphaProof Nexus (8 Oct 2026)](https://www.eurekalert.org/news-releases/1146436) (eurekalert.org, press) 7. [Science Media Centre Spain: expert reactions (8 Oct 2026)](https://sciencemediacentre.es/en/alphaproof-nexus-new-ai-tool-designed-tackle-mathematical-proofs) (sciencemediacentre.es, press) 8. [The Decoder: AlphaProof Nexus solves decades-old math problems for a few hundred dollars (May 2026)](https://the-decoder.com/google-deepminds-alphaproof-nexus-solves-decades-old-math-problems-for-a-few-hundred-dollars/) (the-decoder.com, press) 9. [WION: Google AI solves decades-old maths problems but DeepMind CEO says AGI is still far away (May 2026)](https://embed.wionews.com/technology/google-ai-solves-decades-old-maths-problems-but-deepmind-ceo-says-agi-is-still-far-away-1779695717629) (embed.wionews.com, press) 10. [Pushmeet Kohli on X: thread on DeepMind and mathematics, paper in Science today (8 Oct 2026)](https://x.com/pushmeet/status/2108265638435119494) (x.com, discussion) 11. [Pushmeet Kohli on X: 'Our technical paper on AlphaProof Nexus appears in Science today'](https://x.com/pushmeet/status/2108265646890815730) (x.com, discussion) 12. [Dr Singularity on X: viral summary (10 Oct 2026)](https://x.com/Dr_Singularity/status/2108936442671992990) (x.com, discussion) 13. [erdosproblems.com #12 (DeepMind construction)](https://www.erdosproblems.com/12) (erdosproblems.com, discussion) 14. [erdosproblems.com #846 (DeepMind and OpenAI, independently)](https://www.erdosproblems.com/846) (erdosproblems.com, discussion) 15. [Anatol Wegner: critical analysis of the AlphaProof Nexus paper (31 May 2026)](https://buttondown.com/anatol/archive/deepminds-alphaproof-nexus/) (buttondown.com, discussion) 16. [Hacker News discussion (May 2026)](https://news.ycombinator.com/item?id=48248173) (news.ycombinator.com, discussion) ## Changes - 2026-10-10 (filed): Created (Science publication 8 Oct; the May arXiv preprint had only been mentioned in passing in [OpenAI model disproves Erdős's 80-year-old unit distance conjecture](https://postcutoff.com/e/2026-05-20-ai-disproves-erdos-unit-distance-conjecture/)). The viral X post of 10 Oct was the trigger. ## Related - 2026-10-06: [OpenAI releases 722 AI-written math manuscripts claiming hundreds of open problems](https://postcutoff.com/e/2026-10-06-openai-math-release-722-manuscripts/index.md) - 2026-09-21: [OpenAI says an internal model resolved 100+ long-standing open problems in 24 days of training (list released 6 Oct: 722 manuscripts)](https://postcutoff.com/e/2026-09-21-openai-100-open-problems-claim/index.md) - 2026-09-10: [GPT-6 Astra's Epoch AI run adds more Lean-checked results](https://postcutoff.com/e/2026-09-10-astra-leanopenproblems-september-results/index.md) - 2026-05-20: [OpenAI model disproves Erdős's 80-year-old unit distance conjecture](https://postcutoff.com/e/2026-05-20-ai-disproves-erdos-unit-distance-conjecture/index.md) - 2026-05-09: [Google DeepMind's 'AI co-mathematician' sets FrontierMath Tier 4 record and helps resolve a Kourovka Notebook problem](https://postcutoff.com/e/2026-05-09-deepmind-ai-co-mathematician/index.md) - 2024-07-25: [AlphaProof and AlphaGeometry 2 reach IMO silver-medal standard](https://postcutoff.com/e/2024-07-25-alphaproof-imo-silver/index.md) - People: [Demis Hassabis](https://postcutoff.com/person/demis-hassabis/), [Pushmeet Kohli](https://postcutoff.com/person/pushmeet-kohli/), [George Tsoukalas](https://postcutoff.com/person/george-tsoukalas/), [Swarat Chaudhuri](https://postcutoff.com/person/swarat-chaudhuri/) - Archived: 2 posts and 1 video, listed in https://postcutoff.com/e/2026-10-08-alphaproof-nexus-science-paper/index.json