DeepMind’s New AI Found A Strange New Way To Think
Two Minute PapersYouTube155,477 views as of 10 October 2026
Why it is here
Two Minute Papers on the AlphaProof Nexus preprint (description links arXiv 2605.22763 and google-deepmind/alphaproof-nexus-results); ~155k views, by far the most-watched video on the paper that Science published on 2026-10-08.
Description
Description written by Gemini from the videoGemini 3.8 Flash, 10 October 2026
Summary
Dr. Károly Zsolnai-Fehér of Two Minute Papers reviews Google DeepMind’s paper on AlphaProof Nexus, an AI system that couples large language models with the Lean formal proof verification environment and an Elo-based tournament evaluation loop. The video explains how this tournament harness enables imperfect AI models to autonomously resolve long-standing open mathematical conjectures.
What is shown
- [00:00] Visual animation demonstrating the “AI Proof Tournament System,” matching candidate proofs and updating Elo ratings.
- [00:06] Overview of Hungarian mathematician Paul Erdős and his repository of over 1,000 open mathematical problems.
- [00:21] Visual breakdown of AlphaProof Nexus solving 9 out of 353 attempted Erdős problems with a 97.5% failure rate.
- [00:54] Historical timeline charting AI mathematical reasoning benchmarks from GPT-3 addition errors (4 years ago) to solving 50-year-old open problems today.
- [01:39] Snippets of Lean code formalizations and DeepMind’s preprint “Advancing Mathematics Research with AI-Driven Formal Proof Search”.
- [02:13] Architecture schematic of AlphaProof Nexus, detailing the Mathematician Lean formalization, Prover Subagent, Proof Validator, Rater Subagent (LLM Critic), and Population Database with Elo rankings.
- [03:08] Tree visualization illustrating how candidate branches are scored, pruned, and iterated upon until reaching a validated formal proof.
- [04:05] Conceptual diagram illustrating the shift from increasing raw model size to tightening the multi-agent algorithmic harness around models.
- [05:27] Benchmark table comparing harness configurations on models including Gemini 3.5 Flash, Gemini 3 Flash, and Gemini 3.1 Pro.
- [06:33] Photo of the presenter interviewing Pushmeet Kohli, VP of Science at Google DeepMind.
- [07:03] Sponsor overview of Weights & Biases Weave application tracing and evaluations.
Claims & numbers
- AlphaProof Nexus attempted 353 formalised Erdős problems and successfully proved 9 of them, resulting in a 97.5% failure rate (presenter citing DeepMind).
- The compute cost of the system averaged approximately $200 per problem solved (presenter citing DeepMind).
- DeepMind’s paper also reports resolving 44 out of 492 OEIS (On-Line Encyclopedia of Integer Sequences) conjectures (shown on preprint text at [03:54]).
- Several of the solved conjectures had been open and unsolved for over 50 years (e.g., 56 years).
- Smaller models tested without sufficient scale solved zero problems, demonstrating that capable foundation models remain essential within the harness (presenter claim at [05:18]).
Notable quotes
- [01:34] “Do not look at where we are. Look at where we will be two more papers down the line.”
- [03:49] “A reliable system built out of unreliable parts. I love that.”
- [04:14] “We don’t need to make it smarter, we need to make the harness around it tighter.”
Assessment
This is an independent paper review and educational summary of DeepMind’s AlphaProof Nexus research, supported by motion graphics and diagrams. The animations abstract the iterative search into a stylized tournament visualizer, but the quantitative results and methodology accurately reflect the published DeepMind preprint.
Described by gemini-3.8-flash on 2026-10-10 from the video’s audio and frames.