# AI-driven mathematical & scientific breakthroughs (Post-Cutoff)
Generated 2026-09-29. 119 results. Newest first.

## astronomy

### 2026-03 — RAVEN machine-learning pipeline validates 118 new planets in TESS data
- Problem: Vetting TESS transit candidates at scale
- Result: 118 statistically validated planets and a large vetted candidate catalogue.
- AI system: RAVEN
- Human role: Human-designed pipeline
- Verification: Peer-reviewed in MNRAS; statistical validation
- Status: confirmed
- Sources: [Warwick: AI approach uncovers dozens of hidden planets in TESS data](https://warwick.ac.uk/news/pressreleases/ai-approach-uncovers-dozens-of-hidden-planets/) · [RAVEN TESS paper (arXiv 2603.22597)](https://arxiv.org/abs/2603.22597) · [ScienceDaily: RAVEN validates 118 new planets](https://www.sciencedaily.com/releases/2026/05/260502233926.htm)
### 2025-12 — AI searches 100 million Hubble images in 2.5 days, finding ~1,400 anomalies including 800+ never described
- Problem: Finding rare objects in huge archival datasets
- Result: Hundreds of new anomalous astronomical objects identified in the Hubble archive.
- AI system: AnomalyMatch
- Human role: Human-led with AI tools: AI flagged, humans classified
- Verification: Peer-reviewed in Astronomy & Astrophysics; lens candidates need follow-up
- Status: confirmed
- Sources: [ESA/Hubble: heic2603](https://esahubble.org/news/heic2603/) · [ESA: 1,400 quirky objects found in Hubble's archive](https://www.esa.int/Science_Exploration/Space_Science/1400_quirky_objects_found_in_Hubble_s_archive) · [arXiv 2505.03508](https://arxiv.org/abs/2505.03508)
### 2021-11-22 — NASA's ExoMiner deep-learning model validates 301 new exoplanets from Kepler data
- Problem: Separating real planets from false positives among Kepler transit candidates
- Result: Statistical validation of 301 new exoplanets.
- AI system: ExoMiner
- Human role: Human-designed; outputs reviewed by scientists
- Verification: Peer-reviewed (ApJ); statistical validation, not independent detection
- Status: confirmed
- Sources: [ExoMiner paper (arXiv 2111.10009)](https://arxiv.org/abs/2111.10009) · [NASA JPL: new deep learning method adds 301 planets to Kepler's total count](https://www.jpl.nasa.gov/news/new-deep-learning-method-adds-301-planets-to-keplers-total-count/)
## biology

### 2026-09-23 — Claude agents discover a novel CRISPR-like enzyme system; Anthropic reveals its own biology wet lab [POST-CUTOFF]
- Problem: Discovering new bacterial/phage defence and genome-editing enzyme systems
- Result: Identified 'array-associated reverse transcriptases' (ARTs): phage reverse transcriptases adjacent to long CRISPR-like repeat arrays, from 200,000+ reverse transcriptases and 3,500 candidate systems; biological function still unknown.
- AI system: Claude (≈950 parallel agents)
- Human role: Humans wrote the initial prompt and ran all wet-lab work; agents did the search and analysis
- Verification: Company preprint; experiments ongoing; not peer-reviewed
- Status: disputed
- Why surprising: Scale of the search (950 agents, 21 hours) and an endorsement from CRISPR pioneer Feng Zhang — but critics said finding such clusters is 'the easy part'.
- Sources: [Claude discovers a novel enzyme system with CRISPR-like repeats (Anthropic)](https://www.anthropic.com/news/claude-discovers-novel-enzyme-system) · [Technical preprint (PDF)](https://www-cdn.anthropic.com/22573675ada52a8ca8a97a1a4b4326b2f208a071.pdf) · [TechCrunch: Anthropic says its biology lab has already found something big](https://techcrunch.com/2026/09/23/anthropic-says-its-biology-lab-has-already-found-something-big/) · [TechCrunch: Anthropic is operating a lab that conducts biology experiments](https://techcrunch.com/2026/09/18/anthropic-is-operating-a-lab-that-conducts-biology-experiments/) · [SiliconANGLE: Anthropic opens AI-powered biology research lab](https://siliconangle.com/2026/09/18/anthropic-opens-ai-powered-biology-research-lab/) · [Phys.org: Anthropic touts AI-led biology discovery](https://phys.org/news/2026-09-anthropic-touts-ai-biology-discovery.html) · [MIT Technology Review: When can we say AI made a scientific discovery?](https://www.technologyreview.com/2026/09/28/1145230/when-can-we-say-ai-made-a-scientific-discovery/) · [Irish Times (NYT syndication): Did Anthropic's AI really make a scientific discovery on its own? (Rodríguez Mestre 'jumbotron' priority claim)](https://www.irishtimes.com/world/2026/09/28/did-anthropics-artificial-intelligence-really-make-a-scientific-discovery-on-its-own/) · [Benzinga: scientist says he had already studied the enzymes for 4 years](https://www.benzinga.com/markets/private-markets/26/09/62022648/anthropic-claudes-claimed-breakthrough-in-biology-faces-a-major-question-scientist-says-he-had-already-studied-the-enzymes-for-4-years) · [Inside Anthropic's molecular biology lab (video)](https://www.youtube.com/watch?v=DdCEmlAydcw) · [Anthropic on X: Claude discovers an enzyme system](https://x.com/AnthropicAI/status/2102824959827742916) · [Lucas Harrington on X: genome-mining critique thread](https://x.com/CRISPR_LuCas/status/2102878373160906938)
### 2026-02-10 — Isomorphic Labs unveils IsoDDE drug-discovery engine, hailed as 'an AlphaFold 4' — but proprietary
- Problem: Predicting protein–ligand binding poses and affinities and antibody–antigen structures
- Result: Proprietary engine reported to beat Boltz-2 and physics-based methods on binding-affinity prediction and reach state of the art on antibody–target structures, generalising to molecules unlike its training data.
- AI system: IsoDDE
- Human role: Human-designed system; results reported by the developer
- Verification: Company technical report only; not peer-reviewed, model not released
- Status: pending
- Why surprising: Mohammed AlQuraishi called it 'a major advance, on the scale of an AlphaFold4' while noting 'we know nothing of the details'.
- Sources: [Nature: 'An AlphaFold 4' — scientists marvel at DeepMind drug spin-off's exclusive new AI](https://www.nature.com/articles/d41586-026-00365-7) · [Scientific American (reprint of Nature news)](https://www.scientificamerican.com/article/an-alphafold-4-scientists-marvel-at-deepmind-drug-spin-offs-exclusive-new-ai/)
### 2026-02-05 — GPT-5 autonomously runs 36,000 experiments in Ginkgo's cloud lab, cutting protein-synthesis cost 40%
- Problem: Optimising cell-free protein synthesis cost
- Result: AI-directed autonomous experimentation set a new cost record for cell-free protein production.
- AI system: GPT-5
- Human role: Autonomous experiment design within an automated lab; humans set objective and infrastructure
- Verification: Preprint; lab results; commercial product
- Status: confirmed
- Sources: [OpenAI: GPT-5 lowers protein synthesis cost](https://openai.com/index/gpt-5-lowers-protein-synthesis-cost/) · [bioRxiv preprint](https://www.biorxiv.org/content/10.64898/2026.02.05.703998v1) · [R&D World: GPT-5 autonomously ran 36,000 protein-synthesis experiments](https://www.rdworldonline.com/openais-gpt-5-autonomously-ran-36000-protein-synthesis-experiments-in-ginkgo-bioworks-cloud-lab/)
### 2025-11 — Edison Scientific's Kosmos AI scientist claims six months of research per run
- Problem: Autonomous data analysis and literature synthesis to generate discoveries
- Result: An agent system reproduced three unpublished human findings from raw data and proposed four new findings, with 79.4% of statements judged accurate.
- AI system: Kosmos
- Human role: Humans supply dataset and objective; scientists evaluate the outputs
- Verification: Preprint (arXiv 2511.02824); accuracy assessed by independent scientists hired by the company
- Status: pending
- Sources: [Edison Scientific: Announcing Kosmos](https://edisonscientific.com/news/announcing-kosmos) · [Kosmos: An AI Scientist for Autonomous Discovery (arXiv 2511.02824)](https://arxiv.org/abs/2511.02824) · [Alzforum: Introducing Kosmos, AI scientist makes discoveries overnight](https://www.alzforum.org/news/research-news/introducing-kosmos-ai-scientist-makes-discoveries-overnight)
### 2025-11 — Baker lab designs antibodies from scratch with atomic accuracy using RFdiffusion
- Problem: Designing antibodies to a chosen epitope computationally, without immunisation or library screening
- Result: Epitope-specific antibodies designed de novo, with cryo-EM-validated atomic accuracy.
- AI system: RFdiffusion (antibody-tuned), Chai-2
- Human role: Human-led with AI tools
- Verification: Peer-reviewed in Nature (Baker); preprint (Chai-2); lab-validated
- Status: confirmed
- Sources: [Atomically accurate de novo design of antibodies with RFdiffusion (Nature)](https://www.nature.com/articles/s41586-025-09721-5) · [GeekWire: Nobel winner's lab notches AI-designed antibodies that hit their targets](https://www.geekwire.com/2025/nobel-winners-lab-notches-another-breakthrough-ai-designed-antibodies-that-hit-their-targets/) · [Chai-2 zero-shot antibody design (bioRxiv)](https://www.biorxiv.org/content/10.1101/2025.07.05.663018v1)
### 2025-10-24 — Genentech's GNEprop screens 1.4 billion virtual compounds and finds 82 new antibacterial hits
- Problem: Finding structurally novel antibacterial scaffolds in ultra-large chemical libraries
- Result: Deep-learning virtual screen of 1.4B compounds yielded 82 experimentally confirmed antibacterial hits, ~90x HTS hit rate.
- AI system: GNEprop (graph neural network)
- Human role: Human-led; experimental screening and validation by Genentech scientists
- Verification: Peer-reviewed in Nature Biotechnology; lab-validated in vitro
- Status: confirmed
- Sources: [Nature Biotechnology: Deep-learning-based virtual screening of antibacterial compounds](https://www.nature.com/articles/s41587-025-02814-6) · [Nature Biotechnology commentary: Deep learning speeds the search for new antibiotic scaffolds](https://www.nature.com/articles/s41587-025-02806-6) · [bioRxiv preprint (Sept 2024)](https://www.biorxiv.org/content/10.1101/2024.09.11.612340v1) · [STAT (sponsored): How AI is supercharging antibiotic discovery](https://www.statnews.com/sponsor/2026/01/12/how-ai-is-supercharging-antibiotic-discovery/)
### 2025-09-12 — First AI-generated complete genomes: Evo models design viable bacteriophages that kill resistant E. coli
- Problem: Designing entire functional genomes, not just single genes or proteins
- Result: First viable organisms (bacteriophages) whose complete genomes were designed by a generative AI model.
- AI system: Evo 1, Evo 2
- Human role: Human-led with AI tools: humans set constraints, synthesised and tested genomes
- Verification: Peer-reviewed in Science (2026); lab-validated
- Status: confirmed
- Why surprising: AI designed whole, working genomes of replicating biological entities, some with gene combinations unlike any natural phage.
- Sources: [bioRxiv: generative design of novel bacteriophages with genome language models](https://www.biorxiv.org/content/10.1101/2025.09.12.675911v1) · [Arc Institute: first AI-designed synthetic phage](https://arcinstitute.org/news/hie-king-first-synthetic-phage) · [Stanford News: Evo 2 AI tool designs E. coli-killing bacteriophages (Science, Aug 2026)](https://news.stanford.edu/stories/2026/08/evo-2-ai-tool-e-coli-killer-bacteriophages) · [C&EN: AI program designs new bacteriophages](https://cen.acs.org/biological-chemistry/genomics/ai-program-designs-new-bacteriophages/104/web/2026/08)
### 2025-07 — Stanford's 'Virtual Lab' of AI agents designs SARS-CoV-2 nanobodies validated in the lab
- Problem: Designing nanobodies against newly emerged SARS-CoV-2 variants
- Result: An AI-agent-designed computational pipeline produced nanobodies with improved binding to recent variants.
- AI system: Virtual Lab (GPT-4o agents), ESM, AlphaFold-Multimer, Rosetta
- Human role: AI-assisted: agents designed the workflow; humans gave feedback and ran experiments
- Verification: Peer-reviewed in Nature; lab-validated
- Status: confirmed
- Sources: [The Virtual Lab of AI agents designs new SARS-CoV-2 nanobodies (Nature)](https://www.nature.com/articles/s41586-025-09442-9) · [GitHub: zou-group/virtual-lab](https://github.com/zou-group/virtual-lab)
### 2025-06-25 — AlphaGenome predicts how DNA variants affect thousands of gene-regulation signals from 1 Mb of sequence
- Problem: Predicting the molecular effect of non-coding genetic variants
- Result: Unified sequence-to-function model predicting thousands of regulatory signals and variant effects.
- AI system: AlphaGenome
- Human role: Human-designed model
- Verification: Peer-reviewed in Nature (2026)
- Status: confirmed
- Sources: [DeepMind: AlphaGenome — AI for better understanding the genome](https://deepmind.google/blog/alphagenome-ai-for-better-understanding-the-genome/) · [Nature vol 649 issue 8099 (AlphaGenome paper)](https://www.nature.com/nature/volumes/649/issues/8099) · [Science Media Centre: expert reaction to AlphaGenome](https://www.sciencemediacentre.org/expert-reaction-to-paper-on-google-deepminds-alphagenome/)
### 2025-02-19 — Evo 2: a 40B-parameter genome language model trained on DNA from all domains of life
- Problem: General-purpose modelling and design of DNA across life
- Result: Open genome foundation model enabling zero-shot variant-effect prediction and whole-genome generation.
- AI system: Evo 2
- Human role: Human-designed model
- Verification: Peer-reviewed in Nature (2026)
- Status: confirmed
- Sources: [Arc Institute: Evo 2 one year later](https://arcinstitute.org/news/evo-2-one-year-later) · [Wikipedia: Evo (AI)](https://en.wikipedia.org/wiki/Evo_(AI))
### 2025-02-19 — Google's AI co-scientist independently reproduces an unpublished superbug discovery in 48 hours
- Problem: Mechanism of cf-PICI host-range expansion; drug repurposing for AML and liver fibrosis
- Result: AI-generated hypotheses matching an unpublished discovery and identifying lab-validated drug candidates.
- AI system: AI co-scientist (Gemini 2.0)
- Human role: AI-assisted: scientists pose goals; lab validation by humans
- Verification: Peer-reviewed in Cell (cf-PICI), Advanced Science (fibrosis) and Nature (system, 2026); lab-validated
- Status: confirmed
- Why surprising: Penadés said he initially thought Google had accessed his unpublished data, because the AI's top hypothesis matched his group's years-long result.
- Sources: [Co-Scientist paper (Nature, 2026)](https://www.nature.com/articles/s41586-026-10644-y) · [Cell: AI co-scientist and the cf-PICI mechanism](https://www.cell.com/cell/fulltext/S0092-8674(25)00973-0) · [bioRxiv: cf-PICI hypothesis generated by AI co-scientist](https://www.biorxiv.org/content/10.1101/2025.02.19.639094v1.full) · [Advanced Science: AI-assisted liver fibrosis drug repurposing](https://advanced.onlinelibrary.wiley.com/doi/full/10.1002/advs.202508751) · [HPCwire: Google unveils AI scientist](https://www.hpcwire.com/2025/02/26/google-unveils-ai-scientist-that-could-transform-research/)
### 2024-10-09 — Nobel Prize in Chemistry for protein design and AlphaFold
- Problem: Protein structure prediction and computational protein design
- Result: Nobel Prize in Chemistry 2024: half to David Baker (computational protein design), half to Demis Hassabis and John Jumper (AlphaFold protein structure prediction).
- AI system: AlphaFold 2, Rosetta/RFdiffusion lineage
- Human role: Award recognising human-built AI systems
- Verification: Nobel committee
- Status: confirmed
- Why surprising: First Nobel awarded for an AI system's scientific achievement, only four years after AlphaFold 2.
- Sources: [Nobel Prize in Chemistry 2024 press release (NobelPrize.org)](https://www.nobelprize.org/prizes/chemistry/2024/press-release/) · [Nobel Prize in Chemistry 2024 summary (NobelPrize.org)](https://www.nobelprize.org/prizes/chemistry/2024/summary/) · [Wikipedia: Demis Hassabis](https://en.wikipedia.org/wiki/Demis_Hassabis)
### 2024-09-05 — AlphaProteo designs high-affinity protein binders, including the first AI-designed VEGF-A binder
- Problem: Designing high-affinity binders to disease targets
- Result: Lab-validated de novo binders for 7 targets, including VEGF-A, at high success rates.
- AI system: AlphaProteo
- Human role: Human-designed system; lab testing by collaborators
- Verification: Lab-validated; technical report (not peer-reviewed at release)
- Status: confirmed
- Sources: [DeepMind: AlphaProteo generates novel proteins for biology and health research](https://deepmind.google/blog/alphaproteo-generates-novel-proteins-for-biology-and-health-research/) · [MobiHealthNews: Google DeepMind unveils AlphaProteo](https://www.mobihealthnews.com/news/google-deepmind-unveils-alphaproteo-ai-drug-design)
### 2024-06-25 — ESM3 generates esmGFP, a new fluorescent protein estimated at '500 million years of evolution' from nature
- Problem: Generating functional proteins far from natural sequences
- Result: A functional fluorescent protein with low identity to any natural protein, generated by a language model.
- AI system: ESM3
- Human role: Human-designed prompting and lab validation
- Verification: Peer-reviewed in Science; lab-validated fluorescence
- Status: confirmed
- Sources: [Simulating 500 million years of evolution with a language model (Science)](https://www.science.org/doi/10.1126/science.ads0018) · [EvolutionaryScale: ESM3 release](https://www.evolutionaryscale.ai/blog/esm3-release)
### 2024-05-08 — AlphaFold 3 predicts structures and interactions of all life's molecules
- Problem: Predicting 3D structures of biomolecular complexes (protein–ligand, protein–DNA/RNA, antibodies)
- Result: Single diffusion-based model predicting joint structures of proteins, nucleic acids, ligands and ions, with at least 50% better accuracy on protein–ligand interactions than prior methods.
- AI system: AlphaFold 3
- Human role: Human-designed system; predictions autonomous
- Verification: Peer-reviewed in Nature (May 2024); benchmarked on PoseBusters
- Status: confirmed
- Sources: [Accurate structure prediction of biomolecular interactions with AlphaFold 3 (Nature, DOI)](https://doi.org/10.1038/s41586-024-07487-w) · [AlphaFold 3 predicts the structure and interactions of all of life's molecules (Google)](https://blog.google/technology/ai/google-deepmind-isomorphic-alphafold-3-ai-model/) · [AlphaFold Server](https://alphafoldserver.com/)
### 2023-07-11 — RFdiffusion: diffusion models design new proteins that work in the lab
- Problem: Designing proteins with specified shapes and functions from scratch
- Result: A general generative model whose protein designs fold and bind as intended at high experimental success rates.
- AI system: RFdiffusion
- Human role: Human-designed system; humans select and test designs
- Verification: Peer-reviewed in Nature; lab-validated incl. cryo-EM
- Status: confirmed
- Sources: [De novo design of protein structure and function with RFdiffusion (Nature)](https://www.nature.com/articles/s41586-023-06415-8) · [Baker Lab: RFdiffusion now free and open source](https://www.bakerlab.org/2023/03/30/rf-diffusion-now-free-and-open-source/) · [IPD: RFdiffusion3 now available](https://www.ipd.uw.edu/2025/12/rfdiffusion3-now-available/)
### 2020-11-30 — AlphaFold 2 solves protein structure prediction at CASP14
- Problem: Protein folding / structure prediction problem (open since 1972)
- Result: Median GDT_TS of 92.4 across CASP14 targets — accuracy comparable to experimental structures for most single-chain proteins; later used to predict 200M+ structures.
- AI system: AlphaFold 2
- Human role: Human-designed system; predictions autonomous in blind assessment
- Verification: Blind community assessment (CASP14); peer-reviewed in Nature (2021); widely experimentally corroborated
- Status: confirmed
- Why surprising: CASP co-founder John Moult said the 50-year-old problem had been 'in a sense solved' — years or decades earlier than most structural biologists expected.
- Sources: [AlphaFold: a solution to a 50-year-old grand challenge in biology (DeepMind)](https://deepmind.google/discover/blog/alphafold-a-solution-to-a-50-year-old-grand-challenge-in-biology/) · [Highly accurate protein structure prediction with AlphaFold (Nature, DOI)](https://doi.org/10.1038/s41586-021-03819-2) · [AlphaFold Protein Structure Database](https://alphafold.ebi.ac.uk/)
### 2018-12-02 — AlphaFold (v1) tops the CASP13 protein-structure prediction assessment
- Problem: Protein structure prediction from amino-acid sequence (CASP13 free-modelling targets) (open since 1972)
- Result: Ranked first of ~100 groups at CASP13 by predicting inter-residue distance distributions with a deep network and folding by gradient descent on the resulting potential.
- AI system: AlphaFold 1
- Human role: Human-designed system; predictions made autonomously in a blind assessment
- Verification: Blind community assessment (CASP13); peer-reviewed in Nature (2020)
- Status: confirmed
- Why surprising: A newcomer with no structural-biology track record beat long-established academic groups by a clear margin.
- Sources: [Improved protein structure prediction using potentials from deep learning (Nature, DOI)](https://doi.org/10.1038/s41586-019-1923-7) · [Wikipedia: AlphaFold](https://en.wikipedia.org/wiki/AlphaFold)
## chemistry

### 2023-12-20 — Coscientist: a GPT-4 agent plans and runs real chemistry experiments from plain-English prompts
- Problem: Can an LLM agent autonomously design and execute wet-lab chemistry?
- Result: LLM agent planned and executed real cross-coupling experiments via cloud-lab automation.
- AI system: Coscientist (GPT-4)
- Human role: Autonomous within a supervised lab setup
- Verification: Peer-reviewed in Nature
- Status: confirmed
- Sources: [Autonomous chemical research with large language models (Nature)](https://www.nature.com/articles/s41586-023-06792-0) · [Chemistry World: first GPT-4-powered AI lab assistant](https://www.chemistryworld.com/news/first-gpt-4-powered-ai-lab-assistant-independently-directs-key-organic-reactions/4018723.article)
### 2020-07-08 — Liverpool's mobile robot chemist runs 688 experiments in 8 days and finds a 6× better photocatalyst
- Problem: Optimising photocatalyst formulations for hydrogen production
- Result: Autonomous robotic search found a formulation ~6× more active than the baseline.
- AI system: batched Bayesian optimisation
- Human role: Humans designed the search space and robot workflow; experiments autonomous
- Verification: Peer-reviewed in Nature
- Status: confirmed
- Sources: [A mobile robotic chemist (Nature)](https://www.nature.com/articles/s41586-020-2442-2) · [C&EN: Robot runs almost 700 chemistry experiments](https://cen.acs.org/physical-chemistry/computational-chemistry/Robot-runs-almost-700-chemistry/98/i27)
## climate-weather

### 2026-08-06 — DeepMind open-sources WeatherNext 2 and WeatherNext Cyclones with a Nature paper showing ~1 extra day of hurricane warning [POST-CUTOFF]
- Problem: Forecasting tropical-cyclone track, intensity and size
- Result: WeatherNext Cyclones' 3-day forecasts are about as accurate as prior systems' 2-day forecasts (>24 h extra lead time); weights released for commercial use.
- AI system: WeatherNext Cyclones, WeatherNext 2
- Human role: Human-designed models; used operationally by forecasters at the US National Hurricane Center
- Verification: Peer-reviewed in Nature (2026); operational evaluation with NHC
- Status: confirmed
- Sources: [DeepMind: AI model achieves breakthrough in forecasting cyclones](https://deepmind.google/blog/weathernext-ai-model-achieves-breakthrough-in-forecasting-cyclones/) · [GitHub: google-deepmind/weathernext](https://github.com/google-deepmind/weathernext) · [Google blog: WeatherNext 2 cyclones](https://blog.google/innovation-and-ai/models-and-research/google-deepmind/weathernext-2-cyclones/) · [Open Source For You: DeepMind open sources WeatherNext](https://www.opensourceforu.com/2026/08/google-deepmind-weathernext-ai/)
### 2025-05-21 — Microsoft's Aurora foundation model beats operational forecasts for air quality, waves, cyclones and weather
- Problem: A single model for many environmental forecasting tasks
- Result: Foundation model that, fine-tuned, beats specialised operational systems across several Earth-system tasks.
- AI system: Aurora
- Human role: Human-designed model
- Verification: Peer-reviewed in Nature
- Status: confirmed
- Sources: [A foundation model for the Earth system (Nature)](https://www.nature.com/articles/s41586-025-09005-y) · [Microsoft Source: Aurora goes beyond weather forecasting](https://news.microsoft.com/source/features/ai/microsofts-aurora-ai-foundation-model-goes-beyond-weather-forecasting/)
### 2025-02-25 — AI weather forecasting goes operational: ECMWF's AIFS (Feb 2025), then NOAA's AI models (Dec 2025)
- Problem: Replacing or complementing physics-based operational forecasts
- Result: Machine-learned models adopted as official operational forecasts by major national and international centres.
- AI system: AIFS, AIGFS (GraphCast), AIGEFS, HGEFS
- Human role: Human-built; used by operational forecasters
- Verification: Operational verification by ECMWF and NOAA
- Status: confirmed
- Sources: [ECMWF: AI forecasts become operational](https://www.ecmwf.int/en/about/media-centre/news/2025/ecmwfs-ai-forecasts-become-operational) · [NOAA deploys new generation of AI-driven global weather models](https://www.noaa.gov/news-release/noaa-deploys-new-generation-of-ai-driven-global-weather-models) · [CACM: AI weather forecasting goes operational](https://cacm.acm.org/news/ai-weather-forecasting-goes-operational/)
### 2024-12-04 — GenCast: diffusion-based ensemble forecast beats ECMWF's ENS on 97% of targets
- Problem: Ensemble (probabilistic) medium-range weather forecasting
- Result: First ML ensemble system to outperform the top operational ensemble on the vast majority of targets.
- AI system: GenCast
- Human role: Human-designed model
- Verification: Peer-reviewed in Nature
- Status: confirmed
- Sources: [Probabilistic weather forecasting with machine learning (Nature)](https://www.nature.com/articles/s41586-024-08252-9) · [DeepMind: GenCast](https://deepmind.google/blog/gencast-predicts-weather-and-the-risks-of-extreme-conditions-with-sota-accuracy/)
### 2024-07-22 — NeuralGCM: Google's hybrid physics-ML atmosphere model matches top weather forecasts and runs decades-long climate simulations
- Problem: Fast, accurate general circulation models for both weather and climate
- Result: Hybrid differentiable GCM competitive with ECMWF on medium-range forecasts and able to run decades-long climate simulations at a fraction of the cost.
- AI system: NeuralGCM
- Human role: Human-led research
- Verification: Peer-reviewed in Nature
- Status: confirmed
- Sources: [Nature: Neural general circulation models for weather and climate](https://www.nature.com/articles/s41586-024-07744-y) · [arXiv 2311.07222](https://arxiv.org/abs/2311.07222) · [Google Research: NeuralGCM harnesses AI to better simulate long-range global precipitation](https://research.google/blog/neuralgcm-harnesses-ai-to-better-simulate-long-range-global-precipitation/)
### 2023-11-14 — GraphCast: ML weather model beats the world's best physics-based 10-day forecast on 90% of targets
- Problem: Global medium-range weather prediction
- Result: Learned model outperforming the top operational physics-based deterministic forecast on most variables and lead times.
- AI system: GraphCast
- Human role: Human-designed model; forecasts automated
- Verification: Peer-reviewed in Science; operational adoption
- Status: confirmed
- Why surprising: A model trained on 39 years of reanalysis beat decades of numerical weather prediction engineering on most metrics at a fraction of the compute.
- Sources: [Learning skillful medium-range global weather forecasting (Science)](https://www.science.org/doi/10.1126/science.adi2336) · [DeepMind: GraphCast](https://deepmind.google/blog/graphcast-ai-model-for-faster-and-more-accurate-global-weather-forecasting/) · [NOAA deploys new generation of AI-driven global weather models](https://www.noaa.gov/news-release/noaa-deploys-new-generation-of-ai-driven-global-weather-models)
## computer-science

### 2026-09 — NVIDIA's Nemotron-3-Ultra-CC outscores every human at IOI 2026 (535.4/600) [POST-CUTOFF]
- Problem: International Olympiad in Informatics 2026 problems
- Result: Highest score of any participant, human or AI, on the IOI 2026 problem set.
- AI system: Nemotron-3-Ultra-CC
- Human role: Autonomous during contest
- Verification: Graded by the IOI team per NVIDIA; unofficial entry
- Status: confirmed
- Sources: [NVIDIA AI on X: IOI 2026 result](https://x.com/NVIDIAAI/status/2096032566310789528) · [Post-Training Language Models for Gold-Medal Performance in Coding Competitions (arXiv 2609.02849)](https://arxiv.org/abs/2609.02849) · [AI Weekly: Nvidia's 550B Nemotron beats top human coder at IOI 2026](https://aiweekly.co/alerts/nvidias-550b-nemotron-beats-top-human-coder-at-ioi-2026) · [IOI 2026 statistics](https://stats.ioinformatics.org/olympiads/2026)
### 2026-08-17 — AlphaEvolve helps lower the matrix multiplication exponent ω to below 2.371177 [POST-CUTOFF]
- Problem: Matrix multiplication exponent ω
- Result: New upper bound ω < 2.371177.
- AI system: AlphaEvolve
- Human role: Human-led with AI tools
- Verification: Preprint; bound verifiable from the published optimisation certificates
- Status: confirmed
- Sources: [arXiv 2608.16884](https://arxiv.org/abs/2608.16884) · [AI Weekly: AlphaEvolve helps push matrix multiplication to 2.371177](https://aiweekly.co/alerts/alphaevolve-helps-push-matrix-multiplication-to-2371177) · [Pushmeet Kohli on X announcing ω < 2.371177](https://x.com/pushmeet/status/2089717134129565763)
### 2025-09-27 — Scott Aaronson credits GPT-5 with a key step in a quantum complexity proof
- Problem: Limits of black-box error reduction in QMA
- Result: Proof of tight limits on black-box amplification in QMA, with the central analytic idea proposed by GPT-5.
- AI system: GPT-5-Thinking
- Human role: Human-led with AI tools: humans posed the problem, checked and wrote the proof
- Verification: Expert-checked; arXiv preprint
- Status: confirmed
- Why surprising: A leading complexity theorist said an LLM supplied the idea he would have called 'clever' from a student.
- Sources: [Scott Aaronson: The QMA Singularity](https://scottaaronson.blog/?p=9183) · [Limits to black-box amplification in QMA (arXiv 2509.21131)](https://arxiv.org/abs/2509.21131) · [The Quantum Insider: GPT-5 serves as research assistant](https://thequantuminsider.com/2025/09/29/gpt-5-serves-as-research-assistant-in-proving-one-of-quantum-computing-theorys-trickiest-theorems/)
### 2025-09-22 — AlphaEvolve finds gadgets that prove new NP-hardness of approximation bounds for MAX-k-CUT
- Problem: Inapproximability thresholds for MAX-k-CUT
- Result: New NP-hardness of approximation bounds from AI-discovered gadget reductions.
- AI system: AlphaEvolve
- Human role: Human-led with AI tools: researchers framed the gadget search and proved the theorems
- Verification: Preprint; gadgets verified by exhaustive computation
- Status: confirmed
- Sources: [Reinforced Generation of Combinatorial Structures (arXiv 2509.18057)](https://arxiv.org/abs/2509.18057) · [Google Research: AI as a research partner — advancing theoretical CS with AlphaEvolve](https://research.google/blog/ai-as-a-research-partner-advancing-theoretical-computer-science-with-alphaevolve/)
### 2025-09-17 — AI reaches gold-medal level at the ICPC World Finals
- Problem: ICPC World Finals 2025 problem set (12 problems)
- Result: OpenAI's system solved 12/12 problems (would have ranked 1st); Gemini 2.5 Deep Think solved 10/12, including one no human team solved.
- AI system: OpenAI reasoning models, Gemini 2.5 Deep Think
- Human role: Autonomous under contest time limits
- Verification: Judged by the ICPC official judging system in a supervised setting
- Status: confirmed
- Sources: [Gemini achieves gold-medal level at the ICPC World Finals (Google DeepMind)](https://deepmind.google/blog/gemini-achieves-gold-medal-level-at-the-international-collegiate-programming-contest-world-finals/) · [Wikipedia: International Collegiate Programming Contest](https://en.wikipedia.org/wiki/International_Collegiate_Programming_Contest)
### 2025-05 — Intology's 'Zochi' AI system gets a paper into the ACL 2025 main conference
- Problem: Automated research reaching a top-tier main conference
- Result: An AI-generated paper on automated multi-turn jailbreaking was accepted to the ACL 2025 main conference.
- AI system: Zochi
- Human role: AI-assisted: agent claimed to handle ideation, experiments and writing; humans formatted the manuscript
- Verification: Peer review at ACL 2025 (acceptance confirmed); autonomy self-reported
- Status: confirmed
- Sources: [Intology: Zochi's paper accepted to ACL 2025](https://www.intology.ai/blog/zochi-acl) · [ACL 2025 main conference papers](https://2025.aclweb.org/program/main_papers/) · [LessWrong discussion: Zochi publishes a paper](https://www.lesswrong.com/posts/LtsgfGsXpiLTSGpaW/zochi-publishes-a-paper)
### 2025-03-12 — Sakana's AI Scientist-v2 writes the first fully AI-generated paper to pass peer review (ICLR 2025 workshop)
- Problem: Can an AI system autonomously produce a research paper that passes human peer review?
- Result: A fully AI-generated ML paper passed peer review at an ICLR 2025 workshop (scores 6/7/6), a first for end-to-end AI-authored research.
- AI system: The AI Scientist-v2
- Human role: Humans chose the broad topic and selected 3 of the generated papers to submit; no human edits to the paper
- Verification: Blind peer review at an ICLR workshop; system described in Nature (2026)
- Status: confirmed
- Why surprising: A paper with no human-written text beat more than half of the human submissions in blind review.
- Sources: [Sakana AI: The AI Scientist generates its first peer-reviewed scientific publication](https://sakana.ai/ai-scientist-first-publication/) · [Sakana AI: The AI Scientist published in Nature](https://sakana.ai/ai-scientist-nature/) · [Nature news on the AI Scientist paper](https://www.nature.com/articles/d41586-026-00899-w) · [The AI Scientist (v1) paper, arXiv 2408.06292](https://arxiv.org/abs/2408.06292)
### 2023-06-07 — AlphaDev discovers faster small-sort routines, merged into LLVM's C++ standard library
- Problem: Optimal assembly for fixed-size sorting
- Result: RL-discovered sort3–sort5 routines with fewer instructions than human-written libc++ code, adopted upstream.
- AI system: AlphaDev
- Human role: Autonomous search; humans integrated code into LLVM
- Verification: Peer-reviewed in Nature; code merged into LLVM
- Status: disputed
- Sources: [Faster sorting algorithms discovered using deep reinforcement learning (Nature)](https://www.nature.com/articles/s41586-023-06004-9) · [DeepMind: AlphaDev discovers faster sorting algorithms](https://deepmind.google/blog/alphadev-discovers-faster-sorting-algorithms/) · [Cassio Neri: shorter and faster than Sort3AlphaDev (arXiv 2307.14503)](https://arxiv.org/abs/2307.14503)
### 2022-10-05 — AlphaTensor discovers faster matrix multiplication algorithms, beating Strassen's 1969 record for 4×4 mod 2
- Problem: Minimum number of multiplications for small matrix products (tensor rank) (open since 1969)
- Result: New lower-rank decompositions: 4×4 over GF(2) with 47 multiplications; improvements for several other sizes.
- AI system: AlphaTensor
- Human role: Autonomous search within a human-designed RL game
- Verification: Peer-reviewed in Nature; algorithms checkable by direct computation
- Status: confirmed
- Why surprising: First improvement in over 50 years to a Strassen-era record for a small matrix size.
- Sources: [Discovering faster matrix multiplication algorithms with reinforcement learning (Nature)](https://www.nature.com/articles/s41586-022-05172-4) · [GitHub: google-deepmind/alphatensor](https://github.com/google-deepmind/alphatensor) · [Computational Complexity blog on AlphaTensor](https://blog.computationalcomplexity.org/2022/10/alpha-tensor.html)
## materials

### 2026-09-25 — Lila Sciences' AI-run lab screens 2,942 catalysts and finds iridium- and ruthenium-free palladium oxides for green hydrogen [POST-CUTOFF]
- Problem: Iridium/ruthenium-free anode catalysts for acidic oxygen evolution (PEM water electrolysis)
- Result: AI-guided high-throughput campaign found six Pd-oxide catalyst families; the best is comparable to Ru with 1,000+ h stability.
- AI system: Lila Sciences autonomous lab (Bayesian optimisation + LLMs)
- Human role: AI-directed experiment selection with humans for safety review and partial sample handling
- Verification: Preprint only (arXiv 2609.30133); not peer-reviewed
- Status: pending
- Sources: [Lila: How an AI-run lab cracked open green hydrogen's catalyst problem](https://www.lila.ai/news/how-an-ai-run-lab-cracked-open-green-hydrogens-catalyst-problem) · [arXiv 2609.30133: AI-guided high-throughput discovery of Ir- and Ru-free palladium-oxide catalysts](https://arxiv.org/abs/2609.30133) · [Unite.AI: Lila Sciences' AI lab uncovers palladium catalysts for green hydrogen](https://www.unite.ai/lila-sciences-ai-lab-uncovers-palladium-catalysts-for-green-hydrogen/) · [Bloomberg: Lila Sciences said in talks for funds at $8.5B valuation](https://www.bloomberg.com/news/articles/2026-06-03/lila-sciences-said-in-talks-for-funds-at-8-5-billion-valuation) · [Lila: $350M Series A announcement](https://www.lila.ai/news/announcing-the-close-of-our-series-a) · [MIT Technology Review: AI materials-discovery startups (Dec 2025)](https://www.technologyreview.com/2025/12/15/1129210/ai-materials-science-discovery-startups-investment/)
### 2026-06-29 — Machine-learning screen predicts two new kagome superconductors, confirmed in the lab
- Problem: Predicting new superconductors
- Result: Two superconductors predicted computationally with ML screening and confirmed experimentally.
- AI system: ML pre-screening + quantum geometry theory
- Human role: Human-led with AI tools
- Verification: Peer-reviewed in Physical Review Research; lab-validated
- Status: confirmed
- Sources: [ScienceDaily: Aalto/Rice ML-screened kagome superconductors (Jul 2026)](https://www.sciencedaily.com/releases/2026/07/260701205006.htm) · [Futura Sciences: AI unveils two materials](https://www.futura-sciences.com/en/shock-in-science-ai-unveils-two-materials-that-could-change-everything_39019/)
### 2025-01-16 — Microsoft's MatterGen generates materials to order; flagship result later challenged as a known compound
- Problem: Inverse design of inorganic materials with target properties
- Result: Generative model producing candidate stable materials conditioned on properties; one synthesised with near-target modulus.
- AI system: MatterGen
- Human role: Human-designed model; human synthesis (SIAT/CAS)
- Verification: Peer-reviewed in Nature; one lab synthesis; novelty disputed
- Status: disputed
- Sources: [A generative model for inorganic materials design (Nature)](https://www.nature.com/articles/s41586-025-08628-5) · [Microsoft Research: MatterGen](https://www.microsoft.com/en-us/research/blog/mattergen-a-new-paradigm-of-materials-design-with-generative-ai/) · [whataifound.org: MatterGen finding and critique](https://whataifound.org/finding/2025-01-16-mattergen)
### 2024-01-09 — Microsoft AI and PNNL screen 32 million candidates to find a solid electrolyte using ~70% less lithium
- Problem: Reducing lithium content in solid-state battery electrolytes
- Result: AI-screened new mixed Li/Na solid electrolyte synthesised and demonstrated in a prototype cell.
- AI system: Azure Quantum Elements ML force fields and property models
- Human role: Human-led with AI tools; humans synthesised and tested
- Verification: Lab-synthesised prototype; arXiv preprint
- Status: confirmed
- Sources: [Microsoft Azure blog: how Microsoft's AI screened over 32 million candidates to find a better battery](https://azure.microsoft.com/en-us/blog/quantum/2024/01/09/unlocking-a-new-era-for-scientific-discovery-with-ai-how-microsofts-ai-screened-over-32-million-candidates-to-find-a-better-battery/) · [arXiv 2401.04070](https://arxiv.org/abs/2401.04070) · [Chemistry World: Microsoft's AI system powers new battery discovery](https://www.chemistryworld.com/research/microsofts-ai-and-high-performance-computing-system-powers-new-battery-discovery/4018731.article)
### 2023-11-29 — Berkeley's A-Lab claims 41 new materials from autonomous synthesis; after critiques Nature corrects it to 36 'inorganic' (not 'novel') materials
- Problem: Autonomous robotic synthesis of computationally predicted materials
- Result: Robotic lab synthesised dozens of target inorganic compounds autonomously; the novelty claim was withdrawn after critique.
- AI system: A-Lab (ML-planned synthesis, automated XRD analysis)
- Human role: Autonomous lab operation; human critique and re-analysis
- Verification: Peer-reviewed in Nature; corrected (2026)
- Status: disputed
- Sources: [Nature: A-Lab Author Correction (2026)](https://www.nature.com/articles/s41586-025-09992-y) · [Chemistry World: New analysis raises doubts over autonomous lab's materials discoveries](https://www.chemistryworld.com/news/new-analysis-raises-doubts-over-autonomous-labs-materials-discoveries/4018791.article) · [C&EN: Nature robot chemist paper corrected](https://cen.acs.org/research-integrity/Nature-robot-chemist-paper-corrected/104/web/2026/01)
### 2023-11-29 — GNoME predicts 2.2 million new crystals, 380,000 stable, but novelty and usefulness are disputed
- Problem: Discovering new stable inorganic materials
- Result: An order-of-magnitude expansion of computationally predicted stable crystals (380,000).
- AI system: GNoME
- Human role: Human-designed pipeline; predictions automated
- Verification: Peer-reviewed in Nature; DFT-computed stability; limited experimental synthesis
- Status: disputed
- Why surprising: The scale ('800 years of knowledge') was striking, but so was the pushback from chemists about what counts as a new material.
- Sources: [Scaling deep learning for materials discovery (Nature)](https://www.nature.com/articles/s41586-023-06735-9) · [DeepMind: Millions of new materials discovered with deep learning](https://deepmind.google/blog/millions-of-new-materials-discovered-with-deep-learning/) · [Cheetham & Seshadri critique (Chemistry of Materials)](https://pubs.acs.org/doi/10.1021/acs.chemmater.4c00643)
## mathematics

### 2026-09-21 — OpenAI says an internal model resolved 100+ long-standing open problems in 24 days of training; no list released [POST-CUTOFF]
- Problem: Unspecified 'long-standing open problems'
- Result: Claimed resolution of 100+ open problems; unverified.
- AI system: OpenAI internal model (unnamed)
- Human role: Unknown
- Verification: Unverified claim
- Status: pending
- Sources: [TechCrunch: OpenAI forms math advisory group as its AI resolves more than 100 open problems](https://techcrunch.com/2026/09/21/openai-forms-math-advisory-group-as-its-ai-resolves-more-than-100-open-problems/) · [The Decoder: OpenAI says internal model solved over 100 long-standing math problems](https://the-decoder.com/openai-says-its-internal-model-solved-over-100-long-standing-math-problems-after-just-a-month-of-training/) · [FrontierMath Erdős benchmark (arXiv 2609.25050)](https://arxiv.org/abs/2609.25050) · [OEIS Open benchmark (arXiv 2608.11941)](https://arxiv.org/abs/2608.11941) · [OpenAI: Advisory Group on Mathematics and Artificial Intelligence](https://openai.com/index/advisory-group-on-mathematics-and-ai/) · [Terence Tao blog: Announcing the Advisory Group on Mathematics and Artificial Intelligence](https://terrytao.wordpress.com/2026/09/21/advisory-group-on-mathematics-and-artificial-intelligence/) · [Thomas Bloom on X: FrontierMath Erdős thread](https://x.com/thomasfbloom/status/2095630765035864260)
### 2026-09-11 — Fields Medallists' open letter 'A Severe Misalignment of AI in Mathematics' criticises labs' race for famous problems [POST-CUTOFF]
- Problem: How AI-generated mathematical results should be announced, credited and verified
- Result: Collective statement by Fields Medallists calling out misaligned incentives in AI labs' pursuit of famous problems.
- AI system: n/a
- Human role: Human-led response to AI results
- Verification: Public letter
- Status: confirmed
- Sources: [Terence Tao: A severe misalignment of AI in mathematics](https://terrytao.wordpress.com/2026/09/11/a-severe-misalignment-of-ai-in-mathematics/) · [Scientific American: 25 winners of math's Nobel decry the AI invasion of their discipline](https://www.scientificamerican.com/article/25-winners-of-maths-nobel-prize-decry-the-ai-invasion-of-their-discipline/) · [The crisis of AI-generated mathematics (arXiv 2608.02859)](https://arxiv.org/abs/2608.02859) · [mathandai.org: A Severe Misalignment of AI in Mathematics (declaration text, signatories)](https://mathandai.org/) · [Terence Tao on Mathstodon announcing the declaration](https://mathstodon.xyz/@tao/117253629967855195) · [Timothy Gowers: Why I didn't sign the Fields medallists' letter](https://terrytao.wordpress.com/2026/09/17/why-i-didnt-sign-the-fields-medallists-letter/)
### 2026-09-10 — GPT-6 Astra's Epoch AI run adds more Lean-checked results: Dittert conjecture proved, Ibragimov–Iosifescu and eternal-domination conjectures disproved [POST-CUTOFF]
- Problem: Dittert conjecture; Ibragimov–Iosifescu φ-mixing CLT conjecture; strong n-conjecture (n=4); Gamma–Theta eternal domination conjecture
- Result: One proof (Dittert, all n) and three disproofs, each with a Lean formalization or an explicit checkable counterexample.
- AI system: GPT-6 Astra (pre-release)
- Human role: Autonomous proof search in Epoch AI's harness; Tom Adamczewski directed packaging; Klostermeyer co-wrote the domination paper
- Verification: Formal proofs in Lean (mechanically checked). Statements and write-ups mostly not independently audited
- Status: pending
- Sources: [arXiv 2609.11500: A Counterexample to an Eternal Domination Conjecture](https://arxiv.org/abs/2609.11500) · [GitHub: tadamcz/dittert](https://github.com/tadamcz/dittert) · [GitHub: tadamcz/phi-mixing-clt (Ibragimov–Iosifescu)](https://github.com/tadamcz/phi-mixing-clt) · [GitHub: tadamcz/n-conjecture-strong](https://github.com/tadamcz/n-conjecture-strong) · [VibeMathed: Ibragimov–Iosifescu conjecture status](https://vibemathed.com/problem/ibragimov-iosifescu-varphi-mixing-clt-conjecture) · [Wikipedia: List of mathematical discoveries by artificial intelligence](https://en.wikipedia.org/wiki/List_of_mathematical_discoveries_by_artificial_intelligence)
### 2026-09-08 — OpenAI claims a Millennium Prize problem: 10,000 AI agents prove forced Navier–Stokes blow-up; priority dispute erupts [POST-CUTOFF]
- Problem: Navier–Stokes existence and smoothness (Clay Millennium Prize problem), forced-breakdown case (open since 2000)
- Result: Claimed proof, formalised in Lean, of finite-time blow-up for 3D incompressible Navier–Stokes with smooth external forcing.
- AI system: OpenAI internal model (≈10, 000 parallel agents)
- Human role: Largely autonomous agent swarm; builds on human techniques of Córdoba and Martínez-Zoroa
- Verification: Formal proof in Lean (public); Clay review pending; human peer review ongoing
- Status: disputed
- Why surprising: An AI swarm produced, in under four days, a formally verified proof meeting the letter of a Millennium Prize problem, though experts dispute whether it is the problem that matters.
- Sources: [Sebastien Bubeck on X: allegations are "false and inflammatory"](https://x.com/SebastienBubeck/status/2097214122471432349) · [OfficeChai: Bubeck says he tried to coordinate release with Buckmaster & Alpöge](https://officechai.com/ai/openais-sebastien-bubeck-says-he-tried-to-coordinate-release-of-navier-stokes-related-proofs-with-buckmaster-alpoge-but-was-rebuffed/) · [OpenAI: Navier–Stokes solution](https://openai.com/index/navier-stokes-solution/) · [Quanta: AI has solved one of math's $1 million Millennium Prize problems](https://www.quantamagazine.org/ai-has-solved-one-of-maths-1-million-millennium-prize-problems-20260908/) · [Scientific American: Did OpenAI solve the wrong Navier–Stokes problem?](https://www.scientificamerican.com/article/did-openai-solve-the-wrong-navier-stokes-problem/) · [Terence Tao: finite-time blowup with smooth forcing (Buckmaster–Alpöge–Coiculescu)](https://terrytao.wordpress.com/2026/09/07/finite-time-blowup-with-smooth-forcing-term-for-the-incompressible-porous-medium-boussinesq-and-incompressible-euler-equations/) · [Fortune: OpenAI says it cracked Navier–Stokes; Buckmaster accusation](https://fortune.com/2026/09/08/openai-says-it-cracked-navier-stokes-math-grand-challenge-buckmaster-accusation-cheating-intimidation-tao-lament/) · [CNBC: OpenAI claims to have solved 90-year-old Navier–Stokes problem in 88 hours](https://www.cnbc.com/2026/09/09/openai-navier-stokes-math-problem-solved.html) · [Wikipedia: Navier–Stokes priority controversy](https://en.wikipedia.org/wiki/Navier%E2%80%93Stokes_priority_controversy) · [Scientific American: OpenAI claims blockbuster math breakthrough amid swirl of controversy](https://www.scientificamerican.com/article/openai-claims-blockbuster-math-breakthrough-amid-swirl-of-controversy/) · [Alexander Gamburd: The Siren Call of Silicon Leviathan (arXiv 2609.28591, reflective essay)](https://arxiv.org/abs/2609.28591) · [London Mathematical Society statement on the Navier–Stokes developments (9 Sep)](https://www.lms.ac.uk/news/navier-stokes-equations-breakthrough) · [Anima Anandkumar: Stable singularity of the Euler equations on R³ (concurrent unforced result)](https://terrytao.wordpress.com/2026/09/10/stable-singularity-of-the-euler-equations-on-r3/) · [Techmeme cluster, 2026-09-08](https://www.techmeme.com/260908/p26) · [Startup Fortune: OpenAI recruits nine mathematicians to referee its AI math claims](https://startupfortune.com/openai-recruits-nine-mathematicians-to-referee-its-ais-math-claims/) · [OpenAI on X: Navier-Stokes solution announcement](https://x.com/OpenAI/status/2097374640582668336) · [Noam Brown on X: OpenAI mathematicians' 'Lee Sedol moment'](https://x.com/polynoamial/status/2097375272387613183) · [Tristan Buckmaster on Mastodon: three blow-up results and statement](https://mastodon.social/@tristanbuckmaster/117233413705701198) · [Terence Tao on Mathstodon: Alpöge–Buckmaster, a remarkable achievement](https://mathstodon.xyz/@tao/117233527638291447) · [Terence Tao on Mathstodon: open problems as a non-renewable resource (thread)](https://mathstodon.xyz/@tao/117204929023813310)
### 2026-09-07 — Caltech team (Anandkumar) reports a stable self-similar singularity candidate for the unforced 3D Euler equations on R³, found with PINNs and LLM help [POST-CUTOFF]
- Problem: Finite-time singularity for the unforced 3D incompressible Euler equations on R³ from smooth initial data
- Result: Numerically certified self-similar singular profile plus a (conditional) framework for its nonlinear stability; full rigorous blow-up proof not yet complete.
- AI system: physics-informed neural networks, OpenAI models and other LLMs
- Human role: Human-led; AI (PINNs) discovered the candidate, LLMs simplified bounds and helped formalise derivations
- Verification: Interval-arithmetic certification of the profile; partial Lean formalisation; stability conditional
- Status: pending
- Why surprising: Unlike OpenAI's and Buckmaster–Alpöge's results, it targets the unforced problem on the whole space, which is closer to what physicists care about.
- Sources: [Anima Anandkumar (guest post on Tao's blog): Stable singularity of the Euler equations on R³](https://terrytao.wordpress.com/2026/09/10/stable-singularity-of-the-euler-equations-on-r3/) · [arXiv 2609.10867: Self-Similar Singularity of the Euler Equations on R³](https://arxiv.org/abs/2609.10867) · [arXiv 2609.10860: Stability Framework for the Singularity of the Euler Equations on R³](https://arxiv.org/abs/2609.10860) · [Anandkumar group page on the Euler result](https://tensorlab.cms.caltech.edu/users/anima/euler.html)
### 2026-09-07 — Pre-release GPT-6 Astra disproves the Köthe conjecture (1930) with a Lean-verified counterexample [POST-CUTOFF]
- Problem: Köthe conjecture (open since 1930)
- Result: Explicit counterexample disproving the Köthe conjecture, formally verified in Lean.
- AI system: GPT-6 Astra (pre-release)
- Human role: Autonomous discovery; humans checked and wrote up
- Verification: Formal proof in Lean; pending peer review
- Status: pending
- Why surprising: A 96-year-old central problem of noncommutative ring theory fell as a side effect of a benchmark run.
- Sources: [arXiv 2609.07996 (write-up)](https://arxiv.org/abs/2609.07996) · [GitHub: tadamcz/koethe (Lean proof)](https://github.com/tadamcz/koethe) · [Wikipedia: List of mathematical discoveries by artificial intelligence](https://en.wikipedia.org/wiki/List_of_mathematical_discoveries_by_artificial_intelligence)
### 2026-09-04 — Claude produces the first complete machine-checked proof of Fermat's Last Theorem in Lean, in 11 days [POST-CUTOFF]
- Problem: Formal verification of Fermat's Last Theorem (Wiles 1995)
- Result: First complete machine-checked proof of FLT from the axioms, in Lean 4.
- AI system: Claude (research model comparable to Fable 5.1)
- Human role: Near-autonomous; occasional high-level guidance
- Verification: Formal proof in Lean
- Status: confirmed
- Why surprising: Buzzard's human-led project had expected to need many years to reach a full formalisation; an AI did it in 11 days.
- Sources: [Anthropic: Formalizing Fermat's Last Theorem](https://www.anthropic.com/research/formalizing-fermats-last-theorem) · [AI Weekly: Claude formalized Fermat's Last Theorem in 11 days](https://aiweekly.co/alerts/claude-formalized-fermats-last-theorem-in-11-days-anthropic) · [Anthropic on X: first formalized proof of Fermat's Last Theorem](https://x.com/AnthropicAI/status/2095947707605266436) · [Kevin Buzzard (Xena Project): FLT: Anthropic has beaten me to it](https://xenaproject.wordpress.com/2026/09/04/flt-anthropic-has-beaten-me-to-it/)
### 2026-09-03 — Claude-written Lean proof claims the dying percolation conjecture θ(p_c)=0 in every dimension [POST-CUTOFF]
- Problem: Dying percolation conjecture θ(p_c)=0 for Bernoulli bond percolation on Z^d
- Result: Claimed Lean-verified proof that θ(p_c)=0 for all d ≥ 2, via a new additive gluing inequality that settles Kozma–Nitzan Conjecture 3.
- AI system: Claude (Anthropic), Claude Fable 5.1, GPT-5.6 Sol
- Human role: Autonomous formalization: Claude wrote all the Lean code under Justin Leder's direction. The follow-up proofs of stronger conjectures were produced by GPT-5.6 Sol + Claude Fable 5.1 with 'minimal human intervention' from Ahmed Bou-Rabee
- Verification: Formal proof in Lean (mechanically checked); the statement's fidelity and the informal write-up are not yet refereed
- Status: pending
- Why surprising: One of the central open problems of probability theory, which experts expected to need new ideas, was claimed through a machine-written 87k-line Lean development.
- Sources: [anthropics/formal-math: percolation README (commit 795efb8)](https://github.com/anthropics/formal-math/blob/795efb86f191735c5481675763537cfb4ff37e55/percolation/README.md) · [Gil Kalai: Amazing: There is no Percolation at the Critical Probability in all Dimensions](https://gilkalai.wordpress.com/2026/09/03/amazing-there-is-no-percolation-at-the-critical-probability-in-all-dimensions-solved-by-ai-via-a-conjecture-of-gady-kozma-and-shahaf-nitzan/) · [Ahmed Bou-Rabee: Kozma–Nitzan conjectures verification page](https://nitromannitol.github.io/kn1-verification-b80e9/) · [Kozma & Nitzan: A reduction of the θ(p_c)=0 problem to a conjectured inequality (arXiv 2401.12397)](https://arxiv.org/abs/2401.12397) · [Proofs and Prompts: Applied mathematics has met the machine before (on verification vs validation)](https://proofsandprompts.com/2026/09/28/applied-mathematics-has-met-the-machine-before/) · [Wikipedia: Dying percolation conjecture](https://en.wikipedia.org/wiki/Dying_percolation_conjecture)
### 2026-08-30 — GPT-6 Astra lowers the bounded prime gaps record from 246 to 186 [POST-CUTOFF]
- Problem: Bounded gaps between primes (toward the twin prime conjecture) (open since 2014)
- Result: Claimed proof that infinitely many pairs of primes differ by at most 186.
- AI system: GPT-6 Astra
- Human role: Largely AI-generated per OpenAI
- Verification: Formal proof in Lean (announced); not yet independently peer-reviewed
- Status: pending
- Why surprising: A record that a large Polymath collaboration of top number theorists could not push for 12 years moved by 60 in one AI result.
- Sources: [OpenAI: short gaps between primes (PDF)](https://cdn.openai.com/pdf/51126fac-1b68-4128-9666-c908bcc16033/short_gaps.pdf) · [Julia Stadlmann: Bounded gaps between primes (arXiv 2608.31126; human-only, bound 240)](https://arxiv.org/abs/2608.31126) · [Terence Tao on Mathstodon: Stadlmann shaves 246 to 240 without modern AI tools](https://mathstodon.xyz/@tao/117197525544971208) · [Weijie Su on X (Lean formalisation)](https://x.com/weijie444/status/2095600108956262911)
### 2026-08-26 — GPT-5.6 improves the Erdős–Rankin / Ford–Green–Konyagin–Maynard–Tao bound for large prime gaps [POST-CUTOFF]
- Problem: Large gaps between consecutive primes (Erdős–Rankin; Erdős problem #4)
- Result: Improved lower bound for infinitely many large prime gaps, saving a log log log n factor over FGKMT 2018.
- AI system: GPT-5.6 Pro/Sol, GPT-6 Astra
- Human role: AI-assisted: pseudonymous user DottedCalculator prompted the model; Thomas Bloom wrote the exposition
- Verification: Lean formalization of GPT-6 Astra's version reported; human expert review ongoing
- Status: pending
- Sources: [Erdős problem #4](https://www.erdosproblems.com/4) · [erdosproblems.com forum: problem #4 proof claims](https://www.erdosproblems.com/forum/thread/4/proof-claims) · [Traictory: GPT-5.6 claims a prime-gap record. Who checks the proof?](https://traictory.com/news/2026-09-01-gpt-5-6-prime-gap-math-proofs) · [Wikipedia: List of mathematical discoveries by artificial intelligence](https://en.wikipedia.org/wiki/List_of_mathematical_discoveries_by_artificial_intelligence)
### 2026-08-23 — Claude-assisted search breaks the elliptic curve rank record: rank 30, then 31 [POST-CUTOFF]
- Problem: How large can the rank of an elliptic curve over Q be?
- Result: Explicit elliptic curve with rank at least 31, a new record.
- AI system: Claude
- Human role: AI-assisted search with human researchers (Alpöge, Howell)
- Verification: Independent points checkable by computer; listed by the rank record tables
- Status: confirmed
- Sources: [ICARM: new record-breaking elliptic curve reported](https://icarm.io/news/new-record-breaking-elliptic-curve-reported/) · [Andrej Dujella: history of elliptic curve rank records](https://web.math.pmf.unizg.hr/~duje/tors/rankhist.html) · [Epoch AI open problems: elliptic curve rank](https://epoch.ai/frontiermath/open-problems/elliptic-curve-rank)
### 2026-08-23 — Claude-assisted construction claims a complex structure on the 6-sphere, answering Hopf's 1947 problem (pending verification) [POST-CUTOFF]
- Problem: Hopf problem: does S⁶ admit an integrable complex structure? (open since 1947)
- Result: Claimed explicit construction of integrable complex structures on S⁶ (answer: yes).
- AI system: Claude (internal research model)
- Human role: AI-assisted: human mathematician (Alpöge) directed and wrote up
- Verification: Lean formalisation reported; independent expert verification ongoing
- Status: pending
- Why surprising: If it holds, a 79-year-old problem that resisted generations of geometers, including disputed claimed proofs by Michael Atiyah (2016) and others, was answered with AI help.
- Sources: [Scientific American: AI solves 79-year-old math mystery of six-dimensional spheres](https://www.scientificamerican.com/article/ai-solves-79-year-old-math-mystery-of-six-dimensional-spheres/) · [OfficeChai: Anthropic researcher says Claude helped build a complex structure on S⁶](https://officechai.com/ai/anthropic-researcher-says-claude-helped-build-a-complex-structure-on-s%E2%81%B6-taking-aim-at-the-unsolved-hopf-problem/) · [Follow-up paper (arXiv 2609.26706)](https://arxiv.org/abs/2609.26706)
### 2026-08-18 — Palomar launches: a registry of Lean-verified mathematics to curb misrepresented AI proof claims [POST-CUTOFF]
- Problem: Trustworthy registration of (often AI-generated) formal proofs
- Result: Public registry of Lean-verified results with automated formal and semantic checks.
- AI system: n/a
- Human role: Human-built infrastructure; uses an LLM for semantic-alignment checks
- Verification: Lean Comparator + LLM alignment check
- Status: confirmed
- Sources: [Palomar registry](https://palomar-registry.org/) · [Terence Tao: Palomar, a registry of Lean-verified mathematics](https://terrytao.wordpress.com/2026/08/18/palomar-a-registry-of-lean-verified-mathematics/) · [Palomar statement](https://palomar-registry.org/statement) · [GitHub: leanprover/comparator](https://github.com/leanprover/comparator) · [GitHub: mathlib-initiative/formalization.yaml](https://github.com/mathlib-initiative/formalization.yaml)
### 2026-08-12 — Claude-assisted constructions complete Hadamard matrices for every order below 2000, including 668 [POST-CUTOFF]
- Problem: Hadamard conjecture: a Hadamard matrix exists for every order divisible by 4
- Result: Explicit Hadamard matrices for all 12 previously unknown orders below 2000.
- AI system: Claude
- Human role: AI-assisted search directed by human researchers
- Verification: Computer-checkable constructions
- Status: confirmed
- Sources: [Epoch AI open problems: Hadamard matrix of order 668](https://epoch.ai/frontiermath/open-problems/hadamard) · [John D. Cook: Constructing Hadamard matrices](https://www.johndcook.com/blog/2026/08/13/constructing-hadamard-matrices/)
### 2026-08-10 — Claude proves more than two-thirds of Riemann zeta zeros are simple and on the critical line (up from 41.6%) [POST-CUTOFF]
- Problem: Proportion of nontrivial zeros of ζ(s) on the critical line (toward the Riemann hypothesis)
- Result: Unconditional proof that more than 67.2% of zeta zeros are simple and lie on the critical line.
- AI system: Claude (unreleased research model) in Claude Code
- Human role: Near-autonomous: non-mathematician prompted; experts verified
- Verification: Formal proof in Lean (key results); expert review (Conrey, Goldston); independent re-proof
- Status: confirmed
- Why surprising: A decades-slow line of research toward the Riemann hypothesis jumped by 25 percentage points in one AI run.
- Sources: [anthropics/formal-math: zeta23 Lean formalization](https://github.com/anthropics/formal-math) · [Anthropic: Claude and the zeros of the Riemann zeta function](https://www.anthropic.com/research/riemann-zeta) · [More than two thirds of the zeta zeros are simple and on the critical line (arXiv 2608.13637)](https://arxiv.org/abs/2608.13637) · [Lamzouri: independent proof (arXiv 2609.02882)](https://arxiv.org/abs/2609.02882)
### 2026-08-05 — HRT conjecture (1996) disproved: 12 time-frequency shifts of a Schwartz function are linearly dependent, found with ChatGPT-assisted guesswork [POST-CUTOFF]
- Problem: Heil–Ramanathan–Topiwala (HRT) conjecture on linear independence of time-frequency shifts (open since 1996)
- Result: Explicit counterexample: 12 time-frequency shifts of a Schwartz function are linearly dependent, disproving the HRT conjecture.
- AI system: ChatGPT
- Human role: Human-led with AI assistance for strategy and parameter search; humans wrote and checked the proof
- Verification: Handwritten proof with certified numerics; expert-checked (Tao digest); not formalised in Lean; peer review pending
- Status: confirmed
- Why surprising: A well-known 30-year-old conjecture, widely believed true, turned out false, and the counterexample was found with a chatbot's help.
- Sources: [arXiv 2608.05044: Linear dependence of time-frequency shifts of a Schwartz function](https://arxiv.org/abs/2608.05044) · [Terence Tao: A partial digestion of the HRT counterexample](https://terrytao.wordpress.com/2026/08/06/a-partial-digestion-of-the-hrt-counterexample/)
### 2026-08-05 — Sendov's 1958 conjecture on polynomial roots proved with GPT-5.6 Pro; Tao simplifies and formalises it [POST-CUTOFF]
- Problem: Sendov's conjecture (open since 1958)
- Result: Proof of Sendov's conjecture for all degrees, formally verified in Lean.
- AI system: GPT-5.6 Pro
- Human role: Human orchestrated (Mazur); Tao simplified and formalised with AI agents
- Verification: Formal proof in Lean; expert-checked by Tao
- Status: confirmed
- Why surprising: A 68-year-old conjecture that Tao himself had only proved for large degrees fell to an elementary AI-found argument.
- Sources: [Terence Tao: A digestion of the proof of Sendov's conjecture](https://terrytao.wordpress.com/2026/08/12/a-digestion-of-the-proof-of-sendovs-conjecture/) · [Lech Mazur: Sendov conjecture proof (PDF)](https://www.proofatlas.ai/papers/sendov-conjecture/SENDOV_CONJECTURE_PROOF_AUGUST_5_2026.pdf)
### 2026-08-01 — OpenAI's unreleased 'Astra' model claims ten advances in maths and theoretical CS, with Lean proofs [POST-CUTOFF]
- Problem: Ten open problems incl. existence of explicit non-sofic groups, Connes rigidity, Erdős #146/#180/#183, sphere-packing bounds
- Result: Claimed resolutions or improvements on ten open problems, most with Lean-formalised proofs.
- AI system: Astra (GPT-6 Astra)
- Human role: Largely autonomous per OpenAI; humans selected problems and checked
- Verification: Formal proofs in Lean for most results; independent human audit found no remaining substantive error in principal results
- Status: confirmed
- Why surprising: A single unreleased model produced in one batch results that specialists would count as career highlights, including a question Gromov asked about 25 years earlier.
- Sources: [OpenAI: Ten advances in mathematics and theoretical computer science](https://openai.com/index/ten-advances-in-mathematics/) · [OpenAI: ten proofs manuscript (PDF)](https://cdn.openai.com/pdf/ten-proofs-oai.pdf) · [A Human Audit of OpenAI's AI-Generated Mathematical Proofs (arXiv 2608.14673)](https://arxiv.org/abs/2608.14673) · [Simon Willison on the ten advances](https://simonwillison.net/2026/Aug/1/ten-advances-in-mathematics/) · [Andreas Thom (guest post on Tao's blog): On the existence of non-sofic groups (attribution concerns)](https://terrytao.wordpress.com/2026/09/11/on-the-existence-of-non-sofic-groups/) · [Kun & Thom: Nonsofic wreath products of residually finite groups (arXiv 2608.06222)](https://arxiv.org/abs/2608.06222) · [MathOverflow: key new ideas in the non-soficity proof](https://mathoverflow.net/questions/513866/what-are-the-key-new-ideas-in-the-proof-of-nonsoficity-of-groups-in-openai-s-con) · [Quanta: Why the legendary Erdős problems are falling to AI](https://www.quantamagazine.org/why-the-legendary-erdos-problems-are-falling-to-ai-20260803/)
### 2026-07-27 — Neurosurgery resident uses GPT-5.6 Sol to prove Crouzeix's conjecture in a 16-hour autonomous run [POST-CUTOFF]
- Problem: Crouzeix's conjecture (open since 2004)
- Result: Proof that the numerical range W(A) is a 2-spectral set for every matrix A.
- AI system: GPT-5.6 Sol
- Human role: Autonomous: non-specialist prompted; experts verified
- Verification: Expert-checked (Crouzeix, Greenbaum, Townsend); preprint
- Status: confirmed
- Why surprising: A doctor with no formal maths training settled a well-known conjecture in numerical analysis by letting a model run overnight.
- Sources: [Alex Townsend: SIAM News essay on the Crouzeix conjecture (PDF)](https://alextownsend.net/essays/SIAMNews_CrouzeixConjecture.pdf) · [SCMP: Chinese doctor stuns maths world cracking decades-old problem using ChatGPT](https://www.scmp.com/tech/tech-trends/article/3363966/chinese-doctor-stuns-maths-world-cracking-decades-old-problem-using-chatgpt)
### 2026-07-24 — Terence Tao's ICM 2026 public lecture 'Mathematics in the age of AI' calls a crisis in the foundations of mathematical values [POST-CUTOFF]
- Problem: How the mathematical community should respond to AI tools that can do research-level mathematics
- Result: Programmatic lecture and essay reframing the debate from AI capability to the community's goals and values.
- AI system: n/a
- Human role: Human-led
- Verification: Public lecture; essay submitted to ICM Proceedings
- Status: confirmed
- Sources: [Tao: slides 'Mathematics in the age of AI' (PDF)](https://teorth.github.io/tao-web/slides/age-of-ai-icm-2026.pdf) · [arXiv 2608.16753: Mathematics in the age of AI (essay)](https://arxiv.org/abs/2608.16753) · [Tao on Mathstodon: slides uploaded, AI-made summary and interview](https://mathstodon.xyz/@tao/116977934921819775) · [Terence Tao on AI in mathematics (and beyond), AI-collated summary](https://teorth.github.io/tao-web/ai-views.html) · [Tao: AI 'interview' on his AI views](https://teorth.github.io/tao-web/ai-views-interview.html) · [Scientific American: If AI can do math, what's the point of mathematicians?](https://www.scientificamerican.com/article/mathematicians-confront-the-ai-apocalypse/) · [Simons Foundation: Watch: Terence Tao on AI and why we do math](https://www.simonsfoundation.org/2026/08/13/fields-medalist-terence-tao-on-artificial-intelligence-and-why-we-do-math/) · [YouTube recording (uploaded by Alvaro Lozano-Robledo)](https://www.youtube.com/watch?v=sxAe4HJceFQ)
### 2026-07-24 — Hessian conjecture refuted in five variables, derived from Claude-found Jacobian counterexample [POST-CUTOFF]
- Problem: Hessian conjecture: a polynomial whose Hessian determinant is a nonzero constant has an injective gradient map
- Result: Counterexample for n=5 (hence all n≥5); status now: true for n≤3, false for n≥5, open for n=4
- AI system: Claude Fable 5 (indirectly, via the Jacobian counterexample)
- Human role: Human-led; built by hand on an AI-assisted result
- Verification: arXiv preprint; explicit and checkable
- Status: pending
- Why surprising: An AI-found counterexample spread to a neighbouring conjecture within days.
- Sources: [arXiv 2607.22198: A five-variable counterexample to the Hessian conjecture](https://arxiv.org/abs/2607.22198) · [Terence Tao: A digestion of the Jacobian conjecture counterexample](https://terrytao.wordpress.com/2026/07/21/a-digestion-of-the-jacobian-conjecture-counterexample/)
### 2026-07-23 — AI systems score a perfect 42/42 at IMO 2026, officially graded [POST-CUTOFF]
- Problem: International Mathematical Olympiad 2026 problems (Shanghai)
- Result: First officially graded perfect AI scores at the IMO: Huawei 'Celia' and RedNote 'dots-note-3.0' each solved all six problems for 42/42; several other labs self-reported 42/42.
- AI system: Huawei Celia, RedNote dots-note-3.0
- Human role: Autonomous; problems given after the human contest, no human intervention
- Verification: Graded by IMO organisers (for Celia and dots-note-3.0); other 42/42 claims self-administered
- Status: confirmed
- Why surprising: Only 7 of 666 human contestants got full marks, and the first perfect AI scores came from Huawei and a social-media company rather than a frontier US lab.
- Sources: [TechXplore: AI catches up with humans to score 100% at top math contest](https://techxplore.com/news/2026-07-ai-humans-score-math-contest.html) · [SCMP: RedNote's AI model first to achieve flawless score at maths Olympiad](https://www.scmp.com/tech/article/3361482/worlds-first-ai-model-earn-perfect-score-maths-olympiad-comes-chinas-rednote) · [Taipei Times: AI models score 100 percent at top math competition](https://www.taipeitimes.com/News/world/archives/2026/07/24/2003861308) · [Malay Mail: Huawei, Xiaohongshu AI storm Olympiad](https://www.malaymail.com/news/tech-gadgets/2026/07/23/huawei-xiaohongshu-ai-storm-olympiad-join-maths-elite-with-perfect-100pc-score/228720) · [France 24 / AFP: AI catches up with humans to score 100% at top maths contest](https://www.france24.com/en/live-news/20260723-ai-catches-up-with-humans-to-score-100-at-top-maths-contest) · [Deedy Das on X: self-run IMO 2026 results for frontier models](https://x.com/deedydas/status/2079409461874332066) · [NVIDIA AI on X: Nemotron 3 Ultra graded 30/42 by IMO team](https://x.com/NVIDIAAI/status/2079642933058244704)
### 2026-07-20 — Claude Fable 5 finds a counterexample to the Jacobian conjecture in dimension 3 [POST-CUTOFF]
- Problem: Jacobian conjecture (Keller 1939): a polynomial map with non-zero constant Jacobian determinant is invertible (open since 1939)
- Result: Explicit non-injective polynomial map C³→C³ with constant Jacobian determinant, disproving the conjecture for all n≥3.
- AI system: Claude Fable 5
- Human role: AI-assisted: human chose the problem and verified; Claude found the counterexample
- Verification: Directly computer-checkable; confirmed independently by multiple mathematicians; formal peer review pending
- Status: confirmed
- Why surprising: One of Smale's 18 problems for the 21st century, with a long history of false proofs, was refuted by a short formula found by an AI.
- Sources: [Terence Tao: A digestion of the Jacobian conjecture counterexample](https://terrytao.wordpress.com/2026/07/21/a-digestion-of-the-jacobian-conjecture-counterexample/) · [Shuhong Gao: counterexamples in all dimensions >2 (arXiv 2608.00222)](https://arxiv.org/abs/2608.00222) · [Xena Project: Human mathematicians are being out-counterexampled](https://xenaproject.wordpress.com/2026/07/20/human-mathematicians-are-being-outcounterexampled/) · [ScienceDaily: Claude Fable 5 AI finds a tiny formula that topples an 87-year-old math conjecture](https://www.sciencedaily.com/releases/2026/08/260804034634.htm)
### 2026-07-17 — GPT-5.6 Sol Ultra proves the 50-year-old cycle double cover conjecture [POST-CUTOFF]
- Problem: Cycle double cover conjecture (open since 1973)
- Result: Proof that every bridgeless graph admits a cycle double cover.
- AI system: GPT-5.6 Sol Ultra
- Human role: Autonomous proof per OpenAI; human experts wrote independent expositions
- Verification: Expert-checked (independent expositions); preprint; Lean formalisation reported
- Status: confirmed
- Why surprising: One of graph theory's best-known conjectures, open for about 50 years, was reportedly proved by a model in under an hour.
- Sources: [OpenAI: cycle double cover proof (PDF)](https://cdn.openai.com/pdf/04d1d1e4-bc75-476a-97cf-49055cd98d31/cdc_proof.pdf) · [OpenAI preprint (arXiv 2607.15399)](https://arxiv.org/abs/2607.15399) · [Sang-il Oum: exposition of the proof (arXiv 2607.16356)](https://arxiv.org/abs/2607.16356) · [AI Weekly: OpenAI attributes cycle double cover proof to GPT-5.6 Sol Ultra](https://aiweekly.co/alerts/openai-attributes-cycle-double-cover-proof-to-gpt-56-sol-ultra)
### 2026-07 — AI-assisted counterexample answers Grothendieck's question on finite flat group schemes, merged into Mathlib [POST-CUTOFF]
- Problem: Grothendieck's question: is every finite locally free group scheme of order n killed by n?
- Result: A counterexample of order 4 not killed by 4, over a non-reduced base.
- AI system: GPT-5.6 Sol, Claude Fable 5
- Human role: AI-assisted: Akhil Mathew directed the search and verified
- Verification: Formal proof in Lean (Mathlib)
- Status: confirmed
- Sources: [Benjamin Antieau: Akhil Mathew and AI](https://antieau.github.io/2026/08/10/akhil-mathew-ai.html) · [Xena Project: Human mathematicians are being out-counterexampled](https://xenaproject.wordpress.com/2026/07/20/human-mathematicians-are-being-outcounterexampled/)
### 2026-05-27 — Erdős–Szemerédi sum-product conjecture shown false over the reals; a GPT-5.5 Pro agent re-disproves it in 7 of 8 runs
- Problem: Erdős–Szemerédi sum-product conjecture (over R) (open since 1983)
- Result: Finite sets of reals with both sumset and product set of size at most |A|^(2−c), disproving the conjecture over R; later reproduced autonomously by an AI agent.
- AI system: GPT-5.5 Pro (in the follow-up)
- Human role: Original disproof human-led, inspired by an AI result; follow-up shows autonomous AI rediscovery
- Verification: Human proof (preprint); AI proofs checked by authors
- Status: confirmed
- Sources: [The sum-product conjecture is false for real numbers (arXiv 2605.28781)](https://arxiv.org/abs/2605.28781) · [GPT-5.5 Pro agent disproofs of the sum-product conjecture over R (arXiv 2607.20525)](https://arxiv.org/abs/2607.20525)
### 2026-05-20 — OpenAI model disproves Erdős's 80-year-old unit distance conjecture
- Problem: Erdős unit distance conjecture (planar point sets: at most N^(1+o(1)) unit distances) (open since 1946)
- Result: Construction of planar N-point sets with at least N^(1+δ) unit distances for a fixed tiny δ>0, disproving Erdős's conjectured upper bound, via algebraic number theory; humans improved the exponent within weeks.
- AI system: OpenAI internal reasoning model
- Human role: Autonomous discovery by the model; checked and refined by human mathematicians
- Verification: Expert-checked (Timothy Gowers and others); human follow-up papers
- Status: confirmed
- Why surprising: Widely described as the first historically significant proof produced by an AI; Gowers said he would recommend it to the Annals 'without any hesitation'.
- Sources: [OpenAI: model disproves discrete geometry conjecture](https://openai.com/index/model-disproves-discrete-geometry-conjecture/) · [Human exposition of the counterexample (arXiv 2605.20695)](https://arxiv.org/abs/2605.20695) · [Gil Kalai: Amazing — Erdős unit distance problem was disproved by AI](https://gilkalai.wordpress.com/2026/05/21/amazing-erdos-unit-distance-problem-was-disproved-it-was-achieved-by-ai/) · [Quanta: Why the legendary Erdős problems are falling to AI](https://www.quantamagazine.org/why-the-legendary-erdos-problems-are-falling-to-ai-20260803/) · [Scientific American: AI just solved an 80-year-old Erdős problem](https://www.scientificamerican.com/article/ai-just-solved-an-80-year-old-erdos-problem-and-mathematicians-are-amazed/) · [Physics World: AI-led solutions of Erdős problems spark debate](https://physicsworld.com/a/ai-led-solutions-of-erdos-problems-spark-debate-over-the-future-of-mathematics/) · [MAA: AI solves an 80 year-old Erdős problem](https://maa.org/math-values/ai-solves-an-80-year-old-erdos-problem/) · [Slate: Did A.I. really solve a math problem mathematicians couldn't?](https://slate.com/technology/2026/06/math-chatgpt-erdos-problem-solved-open-ai.html)
### 2026-05-12 — GPT-5.5 Pro finds counterexample disproving McKean's 1966 conjecture and the Gaussian completely monotone conjecture
- Problem: McKean's conjecture on entropy along the heat flow; Gaussian completely monotone (GCM) conjecture (open since 1966)
- Result: Explicit measure on R with positive 5th derivative of entropy under heat flow, refuting three related conjectures.
- AI system: GPT-5.5 Pro
- Human role: AI found the counterexample; humans verified and wrote the proof
- Verification: Human-written rigorous proof; preprint
- Status: confirmed
- Sources: [Gu & Sellke: counterexample to the GCM conjecture (arXiv 2605.11656)](https://arxiv.org/abs/2605.11656) · [Follow-up: multidimensional counterexample (arXiv 2605.18081)](https://arxiv.org/abs/2605.18081) · [Suvrit Sra: GPT, the Counterexample Machine (arXiv 2608.29595)](https://arxiv.org/abs/2608.29595)
### 2026-05-09 — Google DeepMind's 'AI co-mathematician' sets FrontierMath Tier 4 record and helps resolve a Kourovka Notebook problem
- Problem: Kourovka Notebook Problem 21.10; FrontierMath Tier 4
- Result: Resolution of a Kourovka Notebook group-theory problem with a human mathematician; state-of-the-art 48% on FrontierMath Tier 4.
- AI system: AI co-mathematician (Gemini 3.1 Pro)
- Human role: Human-led with AI tools (Lackenby); benchmark autonomous
- Verification: Expert-checked; preprint
- Status: confirmed
- Sources: [AI co-mathematician (arXiv 2605.06651)](https://arxiv.org/abs/2605.06651) · [Epoch AI: new record on FrontierMath Tier 4 (Jan 2026)](https://epochai.substack.com/p/new-record-on-frontiermath-tier-4) · [The Rundown: Google DeepMind's powerful AI co-mathematician](https://www.therundown.ai/p/google-deepmind-powerful-ai-co-mathematician)
### 2026-05-03 — Amateur with GPT-5.4 Pro 'vibe-maths' a 60-year-old Erdős conjecture on primitive sets; Tao co-authors the paper
- Problem: Erdős problem #1196 (Erdős–Sárközy–Szemerédi conjectures on primitive sets) (open since 1966)
- Result: Proof of the conjectured bounds for primitive sets and divisibility chains, plus a new short proof of the Erdős primitive set conjecture.
- AI system: GPT-5.4 Pro
- Human role: AI-generated key idea from a non-expert's prompt; professional mathematicians verified and wrote the paper
- Verification: Expert-checked by Tao, Lichtman and others; arXiv preprint
- Status: confirmed
- Why surprising: A non-mathematician's one-shot prompt produced an elegant proof that a leading expert compared to one 'from The Book'.
- Sources: [Primitive sets and von Mangoldt chains (arXiv 2605.00301)](https://arxiv.org/abs/2605.00301) · [Terence Tao: Primitive sets and von Mangoldt chains — Erdős problem #1196 and beyond](https://terrytao.wordpress.com/2026/05/03/primitive-sets-and-von-mangoldt-chains-erdos-problem-1196-and-beyond/) · [Scientific American: Amateur armed with ChatGPT vibe-maths a 60-year-old problem](https://www.scientificamerican.com/article/amateur-armed-with-chatgpt-vibe-maths-a-60-year-old-problem/)
### 2026-05 — GPT-5.5 Pro-assisted construction lowers the smallest known Borsuk counterexample dimension from 64 to 63
- Problem: Borsuk's conjecture: smallest dimension where it fails (open since 1933)
- Result: Counterexample in dimension 63 (321-point three-distance set needing ≥ 65 parts), improving the 2013 record of 64.
- AI system: GPT-5.5 Pro
- Human role: AI-assisted: Max Grinsztajn worked with GPT-5.5 Pro; exact computer verification
- Verification: Exact finite computation with published certificates; not peer-reviewed
- Status: confirmed
- Why surprising: A 13-year-old record in a famous geometry problem moved by a single added point that a chatbot helped find.
- Sources: [GitHub: maaxgrin/borsuk-63-counterexample (paper PDF + verifier)](https://github.com/maaxgrin/borsuk-63-counterexample) · [Tao et al. optimization constants: constant 28a (Borsuk)](https://teorth.github.io/optimizationproblems/constants/28a.html) · [arXiv 2608.12561: An AI Generated Counterexample to Borsuk Problem in Dimension 63 (withdrawn)](https://arxiv.org/abs/2608.12561) · [Wikipedia: Borsuk's conjecture](https://en.wikipedia.org/wiki/Borsuk%27s_conjecture) · [Wikipedia: List of mathematical discoveries by artificial intelligence](https://en.wikipedia.org/wiki/List_of_mathematical_discoveries_by_artificial_intelligence)
### 2026-03-10 — AlphaEvolve improves lower bounds for nine classical Ramsey numbers
- Problem: Lower bounds for classical two-colour Ramsey numbers R(3,k), R(4,k)
- Result: Explicit colourings improving nine long-studied Ramsey lower bounds.
- AI system: AlphaEvolve
- Human role: Humans set up search and scoring; constructions found by AI
- Verification: Explicit constructions checkable by computer; arXiv preprint
- Status: confirmed
- Sources: [Ramsey lower bounds via AlphaEvolve (arXiv 2603.09172)](https://arxiv.org/abs/2603.09172) · [Wikipedia: Ramsey's theorem (background)](https://en.wikipedia.org/wiki/Ramsey%27s_theorem)
### 2026-03 — Math Inc's Gauss formalises Viazovska's sphere-packing proofs in dimensions 8 and 24, fixing errors in the originals
- Problem: Formal verification of optimal sphere packing in dimensions 8 and 24 (Viazovska 2016; Cohn–Kumar–Miller–Radchenko–Viazovska 2017)
- Result: Complete Lean formalisations of both theorems, correcting minor errors in the published proofs.
- AI system: Gauss
- Human role: AI-assisted: agent built on a human-started blueprint project
- Verification: Formal proof in Lean
- Status: confirmed
- Why surprising: Weeks of agent time formalised a Fields-Medal proof and caught errors that human referees had missed.
- Sources: [Formalizing sphere packing in dimensions 8 and 24 (arXiv 2604.23468)](https://arxiv.org/abs/2604.23468) · [GitHub: math-inc/Sphere-Packing-Lean](https://github.com/math-inc/Sphere-Packing-Lean)
### 2026-02-28 — Donald Knuth's 'Claude's Cycles': Claude Opus 4.6 solves an open Hamiltonian-cycle problem ('Shock! Shock!')
- Problem: Decomposing the 3D torus digraph on m³ vertices into three Hamiltonian cycles
- Result: General explicit construction for all odd m, found by Claude and proved by Knuth.
- AI system: Claude Opus 4.6
- Human role: AI-assisted: colleague Filip Stappers ran the Claude session; Knuth verified and proved
- Verification: Human proof by Knuth; formalised in Lean
- Status: confirmed
- Why surprising: The author of The Art of Computer Programming titled his note's opening 'Shock! Shock!' after an AI solved a problem he had been stuck on.
- Sources: [Donald Knuth: Claude's Cycles (PDF)](https://www-cs-faculty.stanford.edu/~knuth/papers/claude-cycles.pdf) · [GitHub: kim-em/KnuthClaudeLean (Lean formalisation)](https://github.com/kim-em/KnuthClaudeLean) · [Adafruit blog: Don Knuth wrote a paper thanking Claude](https://blog.adafruit.com/2026/03/03/don-knuth-wrote-a-paper-thanking-claude-for-solving-an-open-math-problem/)
### 2026-02-14 — 'First Proof' challenge: AI solves about half of 10 unpublished research problems set by mathematicians
- Problem: Ten previously unpublished lemmas and problems from working mathematicians
- Result: Best AI systems solved roughly 5–6 of 10 fresh research problems, with some disputed gradings and one retraction.
- AI system: Aletheia, OpenAI internal model
- Human role: Autonomous attempts; human expert grading
- Verification: Expert grading by the problem setters
- Status: confirmed
- Sources: [First Proof challenge](https://1stproof.org/) · [OpenAI: First Proof submissions](https://openai.com/index/first-proof-submissions/) · [Scientific American: First Proof is AI's toughest math test yet — the results are mixed](https://www.scientificamerican.com/article/first-proof-is-ais-toughest-math-test-yet-the-results-are-mixed/)
### 2026-02-11 — DeepMind's Aletheia agent and Gemini Deep Think report autonomous Erdős solutions and new physics and CS results
- Problem: Open Erdős problems; conjecture in online optimisation; cosmic-string radiation calculations
- Result: Several Erdős problems solved autonomously, a decade-old online-optimisation conjecture refuted, and a new analytic technique for cosmic-string radiation.
- AI system: Aletheia, Gemini Deep Think
- Human role: Mixed: some autonomous; others in collaboration with 18 external researchers
- Verification: Preprints; expert-checked; one generalisation peer-reviewed
- Status: confirmed
- Sources: [Google DeepMind: Accelerating mathematical and scientific discovery with Gemini Deep Think](https://deepmind.google/blog/accelerating-mathematical-and-scientific-discovery-with-gemini-deep-think/) · [Aletheia paper (arXiv 2602.10177)](https://arxiv.org/abs/2602.10177) · [InfoQ: DeepMind Aletheia agentic math](https://www.infoq.com/news/2026/04/deepmind-aletheia-agentic-math/)
### 2026-01-06 — Erdős problem #728 solved near-autonomously by GPT-5.2 Pro and Harmonic's Aristotle, with a Lean proof
- Problem: Erdős problem #728
- Result: Full solution of the intended version of #728, generated by GPT-5.2 Pro and machine-checked in Lean by Aristotle.
- AI system: GPT-5.2 Pro, Aristotle
- Human role: Near-autonomous: human relayed prompts between two AI systems
- Verification: Formal proof in Lean; endorsed by Terence Tao
- Status: confirmed
- Why surprising: An end-to-end AI pipeline, informal proof plus formal verification, closed an Erdős problem with essentially no human mathematics.
- Sources: [Resolution of Erdős Problem #728: a writeup of Aristotle's Lean proof (arXiv 2601.07421)](https://arxiv.org/abs/2601.07421) · [Terence Tao's wiki: AI contributions to Erdős problems](https://github.com/teorth/erdosproblems/wiki/AI-contributions-to-Erd%C5%91s-problems) · [The Decoder: Tao says GPT-5.2 Pro cracked an Erdős problem but warns the win says more about speed than difficulty](https://the-decoder.com/terence-tao-says-gpt-5-2-pro-cracked-an-erdos-problem-but-warns-the-win-says-more-about-speed-than-difficulty/)
### 2025-12-08 — Genuine AI-assisted solutions to Erdős problems begin: #124 (Aristotle), #1026 (48-hour human–AI collaboration)
- Problem: Erdős problems #124, #367, #707, #1026 and others (open since 1975)
- Result: First AI-involved genuine new solutions of listed-open Erdős problems, including a full solution of #1026 and a Lean-verified autonomous proof for a variant of #124.
- AI system: Aristotle, AlphaEvolve, Gemini Deep Think, GPT-5
- Human role: Mixed: #124 near-autonomous (formal statement given); #1026 human–AI collaboration
- Verification: Formal proofs in Lean for key steps; expert-checked (Tao, Bloom)
- Status: confirmed
- Why surprising: Problems open for decades fell in days, but experts stressed they were obscure ones that few people had seriously attempted.
- Sources: [Terence Tao's wiki: AI contributions to Erdős problems](https://github.com/teorth/erdosproblems/wiki/AI-contributions-to-Erd%C5%91s-problems) · [Terence Tao: The story of Erdős problem #1026](https://terrytao.wordpress.com/2025/12/08/the-story-of-erdos-problem-126/) · [erdosproblems.com forum: problem #124](https://www.erdosproblems.com/forum/thread/124) · [Xena Project: formalization of Erdős problems](https://xenaproject.wordpress.com/2025/12/05/formalization-of-erdos-problems/) · [Alexeev & Mixon on Erdős #707 (arXiv 2510.19804)](https://arxiv.org/abs/2510.19804)
### 2025-12-06 — AxiomProver produces machine-checked Lean proofs for all 12 Putnam 2025 problems
- Problem: William Lowell Putnam Competition 2025 (12 problems)
- Result: Formally verified Lean 4 solutions to all 12 problems, 8 of them within the exam time.
- AI system: AxiomProver
- Human role: Autonomous proof search; humans formalised problem statements (per company)
- Verification: Formal proof in Lean (public repository)
- Status: confirmed
- Why surprising: The hardest undergraduate competition, fully solved with machine-checkable proofs rather than natural-language answers.
- Sources: [GitHub: AxiomMath/putnam2025 (Lean proofs)](https://github.com/AxiomMath/putnam2025) · [Axiom Math: From seeing why to checking everything](https://axiommath.ai/research/from-seeing-why-to-checking-everything/)
### 2025-11-05 — Tao, Gómez-Serrano, Georgiev and Wagner test AlphaEvolve on 67 maths problems
- Problem: Broad battery of optimisation-type open problems (e.g. inequalities, packings, finite-field Kakeya-type constructions)
- Result: Systematic evidence that LLM-driven evolutionary search matches or beats best-known constructions across dozens of problems.
- AI system: AlphaEvolve, Gemini Deep Think, AlphaProof
- Human role: Human-led with AI tools: mathematicians chose problems and scorers
- Verification: Constructions verifiable; preprint
- Status: confirmed
- Sources: [Mathematical exploration and discovery at scale (arXiv 2511.02864)](https://arxiv.org/abs/2511.02864) · [Terence Tao: Mathematical exploration and discovery at scale](https://terrytao.wordpress.com/2025/11/05/mathematical-exploration-and-discovery-at-scale/) · [GitHub: alphaevolve_repository_of_problems](https://github.com/google-deepmind/alphaevolve_repository_of_problems)
### 2025-10-17 — OpenAI researchers claim GPT-5 'solved' 10 Erdős problems; the solutions were already in the literature
- Problem: Problems listed as 'open' on erdosproblems.com
- Result: No new mathematics: GPT-5 located existing published solutions to about 10 listed problems and partial results on 11 others.
- AI system: GPT-5
- Human role: Human researchers prompted literature search and publicised results
- Verification: Checked by Thomas Bloom (site maintainer)
- Status: retracted
- Why surprising: A headline 'breakthrough' collapsed within hours, becoming a cautionary tale about AI maths hype.
- Sources: [TechCrunch: OpenAI's 'embarrassing' math](https://techcrunch.com/2025/10/19/openais-embarrassing-math/) · [The Decoder: OpenAI researcher announced a GPT-5 math breakthrough that never happened](https://the-decoder.com/leading-openai-researcher-announced-a-gpt-5-math-breakthrough-that-never-happened/) · [Terence Tao's wiki: AI contributions to Erdős problems](https://github.com/teorth/erdosproblems/wiki/AI-contributions-to-Erd%C5%91s-problems)
### 2025-09-17 — DeepMind and mathematicians use neural networks to find new unstable singularities in fluid equations
- Problem: Finite-time singularity formation in fluid equations (Euler, Boussinesq, IPM)
- Result: Discovery of previously unknown families of unstable self-similar blow-up solutions, computed to near machine precision.
- AI system: physics-informed neural networks with Gauss–Newton optimisation
- Human role: Human-led with AI tools: co-designed by mathematicians
- Verification: Numerical; preprint; not a rigorous proof
- Status: confirmed
- Sources: [Discovery of unstable singularities (arXiv 2509.14185)](https://arxiv.org/abs/2509.14185) · [Physics World: neural networks discover unstable singularities in fluid systems](https://physicsworld.com/a/neural-networks-discover-unstable-singularities-in-fluid-systems/)
### 2025-09-10 — Math Inc's Gauss agent completes the Strong Prime Number Theorem formalisation in Lean in three weeks
- Problem: Formalising the strong Prime Number Theorem (with error term) in Lean
- Result: Complete machine-checked formalisation produced largely by an AI agent in 3 weeks.
- AI system: Gauss
- Human role: AI-assisted: agent wrote most Lean code from the human blueprint; humans supervised
- Verification: Formal proof in Lean (compiles against Mathlib)
- Status: confirmed
- Why surprising: Weeks of agent time finished a formalisation that expert humans had been working on for a year and a half.
- Sources: [Math Inc: Gauss](https://www.math.inc/gauss) · [GitHub: math-inc/strongpnt](https://github.com/math-inc/strongpnt) · [Math Inc announcement on X](https://x.com/mathematics_inc/status/1966194751847461309)
### 2025-08-20 — GPT-5 Pro proves an improved convex-optimisation bound, which humans had already surpassed
- Problem: Convexity of the gradient-descent optimisation curve for L-smooth convex functions
- Result: A correct, novel proof of the bound η ≤ 1.5/L, improving v1's 1/L but weaker than the humans' tight 1.75/L result already posted.
- AI system: GPT-5 Pro
- Human role: Human posed the problem and verified the proof
- Verification: Expert-checked (Bubeck); informal
- Status: confirmed
- Why surprising: A general chatbot produced a correct, non-trivial research-level proof in minutes, although the result was already superseded.
- Sources: [Sébastien Bubeck on X](https://x.com/SebastienBubeck/status/1958198661139009862) · [whataifound.org: GPT-5 convex bound](https://whataifound.org/finding/2025-08-gpt5-convex-bound) · [What does GPT-5's new math claim actually mean?](https://allthings.how/what-does-gpt-5s-new-math-claim-actually-mean/)
### 2025-07-21 — AI systems reach gold-medal level at the International Mathematical Olympiad
- Problem: International Mathematical Olympiad 2025 problems
- Result: Gemini Deep Think (officially graded) and an experimental OpenAI reasoning model each solved 5 of 6 problems for 35/42 — gold-medal standard — in natural language within the 4.5-hour limits.
- AI system: Gemini Deep Think, OpenAI experimental reasoning model
- Human role: Autonomous during the exam; no human help or formal translation
- Verification: Gemini: certified by IMO coordinators; OpenAI: graded by three former IMO medallists (not officially coordinated)
- Status: confirmed
- Why surprising: In a 2021 public bet Paul Christiano put AI IMO gold by 2025 at under 10% and Eliezer Yudkowsky at about 16%; general-purpose LLMs did it without any formal tools.
- Sources: [Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the IMO (Google DeepMind)](https://deepmind.google/blog/advanced-version-of-gemini-with-deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiad/) · [OpenAI announcement on X](https://x.com/OpenAI/status/1946594928945148246) · [OpenAI Model Earns Gold-Medal Score at International Math Olympiad (Scientific American)](https://www.scientificamerican.com/article/openai-model-earns-gold-medal-score-at-international-math-olympiad-and/) · [Harmonic Aristotle IMO 2025 paper (arXiv 2510.01346)](https://arxiv.org/abs/2510.01346) · [ByteDance Seed-Prover IMO 2025 result](https://seed.bytedance.com/en/blog/bytedance-seed-prover-achieves-silver-medal-score-in-imo-2025)
### 2025-05-14 — AlphaEvolve: Gemini-powered agent discovers new algorithms
- Problem: 4×4 complex matrix multiplication (Strassen 1969: 49 multiplications); kissing number in 11 dimensions; ~50 open problems (open since 1969)
- Result: 48 scalar multiplications for 4×4 complex matrices; kissing configuration of 593 spheres in 11D (previous 592); matched SOTA on ~75% and improved ~20% of 50+ problems.
- AI system: AlphaEvolve, Gemini 2.0 Flash, Gemini 2.0 Pro
- Human role: Humans define the problem and an automated scorer; evolutionary LLM search autonomous
- Verification: Constructions verified computationally (independent GitHub checks); white paper, later arXiv
- Status: confirmed
- Why surprising: The first improvement on Strassen's 4×4 complex case in 56 years came from a general-purpose coding agent, not a specialised system like AlphaTensor.
- Sources: [AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms (Google DeepMind)](https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/) · [AlphaEvolve: A coding agent for scientific and algorithmic discovery (arXiv)](https://arxiv.org/abs/2506.13131) · [Independent verification of the 48-multiplication algorithm (GitHub)](https://github.com/PhialsBasement/AlphaEvolve-MatrixMul-Verification) · [Human follow-up: 48 multiplications with rational coefficients (arXiv 2506.13242)](https://arxiv.org/abs/2506.13242)
### 2024-12-20 — OpenAI o3 scores 25% on FrontierMath research-level maths benchmark, amid funding disclosure controversy
- Problem: FrontierMath: unpublished research-level problems with automatically checkable answers
- Result: Score jump from <2% to a claimed 25.2% in about six weeks, later partly qualified by independent evaluation.
- AI system: OpenAI o3
- Human role: Autonomous answering; benchmark written by expert mathematicians
- Verification: Company-reported; independent Epoch evaluation of released o3 was lower
- Status: disputed
- Why surprising: When FrontierMath launched, Fields medallists including Terence Tao said its problems would likely resist AI for years; a big jump came within weeks.
- Sources: [Epoch AI: OpenAI and FrontierMath](https://epoch.ai/latest/openai-and-frontiermath) · [The Decoder: OpenAI quietly funded independent math benchmark](https://the-decoder.com/openai-quietly-funded-independent-math-benchmark-before-setting-record-with-o3/) · [TechRepublic: independent FrontierMath score for o3](https://www.techrepublic.com/article/news-openai-generative-ai-models-frontiermath-score/)
### 2024-07-25 — AlphaProof and AlphaGeometry 2 reach IMO silver-medal standard
- Problem: International Mathematical Olympiad 2024 problems
- Result: Solved 4 of 6 IMO 2024 problems (28/42, one point below gold) with machine-checked Lean proofs (AlphaProof) and AlphaGeometry 2, including the hardest problem (P6).
- AI system: AlphaProof, AlphaGeometry 2
- Human role: Humans translated problems into Lean; proofs found autonomously (up to 3 days of compute)
- Verification: Formal proof in Lean; graded by IMO medalists Timothy Gowers and Joseph Myers
- Status: confirmed
- Why surprising: Fields medallist Timothy Gowers, who graded the solutions, publicly described the result as well beyond what he had thought was the state of the art in automated theorem proving.
- Sources: [AI achieves silver-medal standard solving International Mathematical Olympiad problems (Google DeepMind)](https://deepmind.google/discover/blog/ai-solves-imo-problems-at-silver-medal-level/) · [AlphaGeometry: Solving olympiad geometry without human demonstrations (Nature, DOI)](https://doi.org/10.1038/s41586-023-06747-5) · [Olympiad-level formal mathematical reasoning with reinforcement learning (AlphaProof, Nature 2025)](https://www.nature.com/articles/s41586-025-09833-y)
### 2024-01-17 — AlphaGeometry solves olympiad geometry near gold-medallist level without human demonstrations
- Problem: Olympiad-level geometry proofs
- Result: 25 of 30 IMO geometry problems solved with human-readable proofs.
- AI system: AlphaGeometry
- Human role: Autonomous
- Verification: Peer-reviewed in Nature; proofs checked by an IMO coach (Evan Chen)
- Status: confirmed
- Sources: [Solving olympiad geometry without human demonstrations (Nature)](https://www.nature.com/articles/s41586-023-06747-5) · [Nature news on AlphaGeometry](https://www.nature.com/articles/d41586-024-00145-1) · [Wu's method can boost symbolic AI to rival silver medalists (arXiv 2404.06405)](https://arxiv.org/abs/2404.06405)
### 2023-12-14 — FunSearch: an LLM finds new cap-set constructions, the first LLM discovery in open maths
- Problem: Cap set problem: largest subsets of F_3^n with no three points on a line
- Result: Largest known cap set in dimension 8 (512 vs 496) and improved capacity lower bound.
- AI system: FunSearch (PaLM 2 / Codey)
- Human role: Humans wrote program skeleton and evaluator; LLM-driven search autonomous
- Verification: Peer-reviewed in Nature; constructions verified computationally
- Status: confirmed
- Why surprising: First time an LLM-based system produced a verifiably new result on a well-known open combinatorics problem.
- Sources: [Mathematical discoveries from program search with large language models (Nature)](https://www.nature.com/articles/s41586-023-06924-6) · [GitHub: google-deepmind/funsearch](https://github.com/google-deepmind/funsearch) · [Ernest Davis: comment on FunSearch](https://cs.nyu.edu/~davise/papers/FunSearchComment.pdf)
### 2021-12-01 — DeepMind and mathematicians use machine learning to guide new theorems in knot theory and representation theory
- Problem: Combinatorial invariance conjecture for Kazhdan–Lusztig polynomials; relations between knot invariants
- Result: ML-guided discovery of a conjectured, then proved, relation between the knot signature and hyperbolic invariants, and a new approach to combinatorial invariance for symmetric groups.
- AI system: supervised neural networks with gradient saliency
- Human role: Human-led with AI tools: ML suggested patterns; mathematicians formulated and proved the theorems
- Verification: Peer-reviewed in Nature; human proofs
- Status: confirmed
- Sources: [Advancing mathematics by guiding human intuition with AI (Nature)](https://www.nature.com/articles/s41586-021-04086-x) · [Critical review of the paper (arXiv 2112.04324)](https://arxiv.org/abs/2112.04324)
### 2021-04-29 — Adam Zsolt Wagner uses reinforcement learning to find counterexamples to open graph-theory conjectures
- Problem: Several published conjectures on graph invariants
- Result: Explicit counterexamples found by an RL agent that treats building a graph as a game, rewarded by how badly the conjecture fails.
- AI system: deep cross-entropy RL
- Human role: Human chose conjectures and reward functions; search autonomous; counterexamples trivially checkable
- Verification: Counterexamples checkable by direct computation; reimplemented by others (arXiv 2403.18429)
- Status: confirmed
- Sources: [Constructions in combinatorics via neural networks (arXiv 2104.14516)](https://arxiv.org/abs/2104.14516) · [Reimplementation and extension (arXiv 2403.18429)](https://arxiv.org/abs/2403.18429)
## medicine

### 2026-09-14 — FDA grants priority review to Takeda's zasocitinib, a computationally designed TYK2 inhibitor, with a decision due Q1 2027 [POST-CUTOFF]
- Problem: Selective allosteric TYK2 inhibition for psoriasis
- Result: Computationally designed oral TYK2 inhibitor met all Phase 3 endpoints, beat deucravacitinib head-to-head and entered FDA priority review.
- AI system: Schrödinger FEP+ physics-based platform, machine learning
- Human role: Human-led medicinal chemistry with physics-based computation and ML
- Verification: Phase 3 randomized trials; FDA review pending
- Status: pending
- Sources: [Takeda: FDA accepts zasocitinib NDA with priority review](https://www.takeda.com/newsroom/newsreleases/2026/fda-priority-review-zasocitinib-psoriasis/) · [PharmaVoice: Nimbus used AI to help develop Takeda's $4B psoriasis bet](https://www.pharmavoice.com/news/nimbus-takeda-zasocitinib-ai-drug-discovery/831289/) · [BioSpace: Takeda's $4B Nimbus bet pays off with best-in-class Phase III data](https://www.biospace.com/drug-development/takedas-4b-nimbus-bet-pays-off-with-best-in-class-phase-iii-plaque-psoriasis-data) · [IntuitionLabs: AI drug discovery FDA approvals, 2026 reality check](https://intuitionlabs.ai/articles/ai-drug-discovery-fda-approvals)
### 2026-09-10 — First Phase III trial of a generative-AI-discovered drug doses first patient (Insilico's rentosertib) [POST-CUTOFF]
- Problem: Idiopathic pulmonary fibrosis (IPF) therapy via a novel target
- Result: AI-identified target (TNIK) and AI-designed molecule (rentosertib) reached Phase III after a Phase IIa in Nature Medicine showing +98.4 mL mean FVC at 12 weeks on 60 mg vs a decline on placebo.
- AI system: Insilico Pharma.AI (PandaOmics, Chemistry42)
- Human role: AI-assisted: AI proposed target and molecule; human chemists, clinicians and regulators ran development and trials
- Verification: Phase IIa peer-reviewed in Nature Medicine (June 2025); Phase III ongoing
- Status: pending
- Why surprising: The first drug with both target and molecule discovered by generative AI entered Phase III about five years after target discovery.
- Sources: [Insilico: first patient dosed in GENESIS-IPF-3](https://insilico.com/news/isn1009261-insilico-medicine-doses-first-patient-genesis-ipf-3) · [PR Newswire: Insilico initiates Phase III trial for rentosertib](https://www.prnewswire.com/news-releases/insilico-initiates-phase-iii-clinical-trial-for-rentosertib-its-ai-empowered-tnik-inhibitor-for-idiopathic-pulmonary-fibrosis-302819553.html) · [EurekAlert: Nature Medicine publishes rentosertib Phase IIa results (June 2025)](https://www.eurekalert.org/news-releases/1086096) · [Drug Target Review: Insilico begins Phase III of AI-designed drug](https://www.drugtargetreview.com/insilico-medicine-launches-phase-iii-trial-of-ai-designed-rentosertib-drug/2135890.article)
### 2026-09-08 — AlphaGenome Atlas predicts the effect of all ~9 billion possible single-letter human DNA variants [POST-CUTOFF]
- Problem: Diagnosing rare diseases caused by non-coding variants
- Result: Genome-wide variant-effect atlas with claimed doubling of rare-disease variant identification.
- AI system: AlphaGenome
- Human role: Human-designed; clinical collaborators validated cases
- Verification: Technical paper; company-reported benchmarks; some cases lab-validated
- Status: pending
- Sources: [Fortune: Google DeepMind AI predictions for 9 billion mutations in the human genome](https://fortune.com/2026/09/08/google-deepmind-ai-predictions-9-billion-mutation-human-genome/) · [DeepMind: AlphaGenome](https://deepmind.google/blog/alphagenome-ai-for-better-understanding-the-genome/)
### 2026-06-05 — Computationally designed broad coronavirus vaccine is safe and immunogenic in first human trial
- Problem: Broadly protective vaccines against current and future coronaviruses
- Result: A computationally designed pan-sarbecovirus antigen was safe and broadly immunogenic in humans.
- AI system: DIOSynVax computational antigen design platform
- Human role: Human-led with computational/AI tools
- Verification: Phase 1 clinical trial; peer-reviewed
- Status: confirmed
- Sources: [ScienceDaily: computer-designed coronavirus vaccine tested in people](https://www.sciencedaily.com/releases/2026/06/260605023357.htm) · [DIOSynVax](https://www.diosynvax.com/)
### 2025-10-15 — Google's C2S-Scale 27B model generates a new cancer-immunotherapy hypothesis confirmed in living cells
- Problem: Making 'cold' tumours visible to the immune system
- Result: Novel, context-dependent drug synergy predicted by an LLM-style single-cell model and confirmed in cell experiments.
- AI system: C2S-Scale 27B (Gemma)
- Human role: AI-generated hypothesis; human lab validation
- Verification: Lab-validated in vitro; preprint
- Status: confirmed
- Sources: [Google: How a Gemma model helped discover a new potential cancer therapy pathway](https://blog.google/technology/ai/google-gemma-ai-cancer-therapy-discovery/) · [DDW: Google AI model reveals new way to improve immunotherapy](https://www.ddw-online.com/google-ai-model-reveals-new-way-to-improve-immunotherapy-38114-202510/)
### 2025-08-14 — Generative AI designs new antibiotics that kill drug-resistant gonorrhoea and MRSA
- Problem: Designing entirely new antibiotic chemotypes against resistant bacteria
- Result: De novo AI-generated molecules with novel mechanisms, effective against drug-resistant gonorrhoea (in vitro) and MRSA (in mice).
- AI system: generative chemistry models (CReM, F-VAE) with GNN property predictors
- Human role: Human-led with AI tools: humans synthesised and tested
- Verification: Peer-reviewed in Cell; lab-validated
- Status: confirmed
- Why surprising: The molecules were not found in any library but invented by AI, and they work through mechanisms not seen in existing drugs.
- Sources: [MIT News: Using generative AI, researchers design compounds that can kill drug-resistant bacteria](https://news.mit.edu/2025/using-generative-ai-researchers-design-compounds-kill-drug-resistant-bacteria-0814) · [Euronews: MIT scientists use AI to develop new antibiotics for gonorrhoea and MRSA](https://www.euronews.com/health/2025/08/15/mit-scientists-use-ai-to-develop-new-antibiotics-for-stubborn-gonorrhoea-and-mrsa)
### 2025-05-20 — FutureHouse's Robin multi-agent system proposes ripasudil as a new treatment candidate for dry AMD
- Problem: Treatments for dry age-related macular degeneration
- Result: AI-generated hypothesis that ripasudil enhances RPE phagocytosis, confirmed in cell assays.
- AI system: Robin (Crow, Falcon, Finch agents)
- Human role: AI-assisted: AI generated hypotheses and analyses; humans ran experiments
- Verification: Peer-reviewed in Nature (2026); lab-validated in vitro
- Status: confirmed
- Sources: [Robin paper (Nature, 2026)](https://www.nature.com/articles/s41586-026-10652-y) · [FutureHouse: Demonstrating end-to-end scientific discovery with Robin](https://www.futurehouse.org/research-announcements/demonstrating-end-to-end-scientific-discovery-with-robin-a-multi-agent-system)
### 2025-01-15 — AI-designed proteins neutralise deadly snake-venom toxins and protect mice
- Problem: Neutralising snake-venom three-finger toxins, poorly handled by existing antivenoms
- Result: De novo designed proteins that neutralise lethal toxins in mice.
- AI system: RFdiffusion, ProteinMPNN
- Human role: Human-led with AI tools
- Verification: Peer-reviewed in Nature; lab-validated in mice
- Status: confirmed
- Sources: [De novo designed proteins neutralize lethal snake venom toxins (Nature)](https://www.nature.com/articles/s41586-024-08393-x) · [Baker Lab: Neutralizing deadly snake toxins](https://www.bakerlab.org/2025/01/15/neutralizing-deadly-snake-toxins/) · [DTU: AI-designed proteins neutralise snake toxins](https://www.dtu.dk/english/newsarchive/2025/01/ai-designed-proteins-neutralise-snake-toxins)
### 2023-12-20 — Explainable deep learning discovers a new structural class of antibiotics against MRSA
- Problem: New antibiotic classes against MRSA
- Result: First new structural class of antibiotics found via explainable deep learning, validated in mouse infection models.
- AI system: graph neural networks with Monte Carlo tree search rationale extraction
- Human role: Human-led with AI tools
- Verification: Peer-reviewed in Nature; lab-validated in mice
- Status: confirmed
- Sources: [Discovery of a structural class of antibiotics with explainable deep learning (Nature)](https://www.nature.com/articles/s41586-023-06887-8) · [Broad Institute: Researchers use AI to identify new class of antibiotic candidates](https://www.broadinstitute.org/news/researchers-use-ai-identify-new-class-antibiotic-candidates)
### 2023-09-19 — AlphaMissense classifies 89% of all 71 million possible human missense mutations
- Problem: Interpreting 'variants of uncertain significance' in rare-disease diagnosis
- Result: Proteome-wide pathogenicity predictions for every possible missense variant.
- AI system: AlphaMissense
- Human role: Human-designed; predictions automated
- Verification: Peer-reviewed in Science; benchmarked against clinical databases
- Status: confirmed
- Sources: [Accurate proteome-wide missense variant effect prediction with AlphaMissense (Science)](https://www.science.org/doi/10.1126/science.adg7492) · [DeepMind: A catalogue of genetic mutations to help pinpoint the cause of diseases](https://deepmind.google/blog/a-catalogue-of-genetic-mutations-to-help-pinpoint-the-cause-of-diseases/)
### 2023-05-25 — AI finds abaucin, a narrow-spectrum antibiotic against the superbug Acinetobacter baumannii
- Problem: Drugs for WHO-priority pathogen A. baumannii
- Result: A new narrow-spectrum antibiotic with a novel mechanism, effective in a mouse wound infection model.
- AI system: graph neural network (Chemprop)
- Human role: Human-led with AI tools
- Verification: Peer-reviewed in Nature Chemical Biology; lab-validated in mice
- Status: confirmed
- Sources: [Deep learning-guided discovery of an antibiotic targeting Acinetobacter baumannii (Nat Chem Biol)](https://www.nature.com/articles/s41589-023-01349-8) · [MIT News: Using AI, scientists find a drug that could combat drug-resistant infections](https://news.mit.edu/2023/using-ai-scientists-combat-drug-resistant-infections-0525)
### 2020-02-20 — Deep learning discovers halicin, a structurally new broad-spectrum antibiotic
- Problem: Finding new antibiotic classes against drug-resistant bacteria
- Result: Discovery of halicin, a broad-spectrum bactericidal compound unlike existing antibiotics, validated in mice.
- AI system: Chemprop message-passing neural network
- Human role: Human-led with AI tools: humans built training data and did all lab validation
- Verification: Peer-reviewed in Cell; lab-validated in vitro and in mice
- Status: confirmed
- Why surprising: The first time deep learning found a new antibiotic from scratch, among molecules that chemists had not considered antibacterial.
- Sources: [A Deep Learning Approach to Antibiotic Discovery (Cell)](https://www.cell.com/cell/fulltext/S0092-8674(20)30102-1) · [PubMed record](https://pubmed.ncbi.nlm.nih.gov/32084340/) · [Chemistry World: AI tool screens 107 million molecules, discovers potent new antibiotics](https://www.chemistryworld.com/news/ai-tool-screens-107-million-molecules-discovers-potent-new-antibiotics/4011233.article)
## other

### 2025-11-20 — OpenAI publishes 'Early science acceleration experiments with GPT-5', including four new math results
- Problem: Can a frontier LLM contribute to research across disciplines?
- Result: Documented cases of GPT-5 contributing proofs, literature finds and hypotheses, including four new maths results that the human authors verified.
- AI system: GPT-5, GPT-5 Pro
- Human role: Human-led with AI tools: experts posed problems, steered the model and verified all outputs
- Verification: Expert-checked by the named co-authors; preprint, not peer-reviewed
- Status: confirmed
- Sources: [Early science acceleration experiments with GPT-5 (arXiv 2511.16072)](https://arxiv.org/abs/2511.16072) · [OpenAI: Accelerating science with GPT-5](https://openai.com/index/accelerating-science-gpt-5/) · [Alex Lupsasca on GPT-5 Pro and black-hole symmetries (OpenAI Academy)](https://academy.openai.com/public/blogs/alex-lupsasca-gpt-5-pro-black-hole-physics-hidden-symmetries) · [OpenAI: GPT-5 and an immunology mystery](https://openai.com/index/gpt-5-immunology-mystery/)
### 2025-10-22 — Agents4Science 2025: first conference where AI must be first author and reviewer
- Problem: How good is AI-authored and AI-reviewed science?
- Result: A full conference cycle with AI first authors and LLM reviewers: 48 of 315 submissions accepted; organisers published an analysis of AI reviewer behaviour.
- AI system: GPT-5, Gemini 2.5, Claude Sonnet 4, various author agents
- Human role: Humans organised, co-authored and spot-checked reviews
- Verification: Conference proceedings and analysis paper (arXiv 2511.15534)
- Status: confirmed
- Sources: [Agents4Science analysis paper (arXiv 2511.15534)](https://arxiv.org/abs/2511.15534) · [Agents4Science accepted papers](https://agents4science.stanford.edu/accepted-papers.html) · [Nature news on the AI-authored conference](https://www.nature.com/articles/d41586-025-03363-3) · [Science News: a science conference tests AI agents](https://www.sciencenews.org/article/science-conference-test-ai-agents)
## physics

### 2026-09-03 — PPPL's PACMAN framework lets multiple AI models control a tokamak in ~20 ms, preventing a tearing mode [POST-CUTOFF]
- Problem: Integrating multiple AI predictors and controllers safely into real-time fusion operation
- Result: Modular real-time AI control framework demonstrated on DIII-D across several control tasks.
- AI system: PACMAN framework (RL and predictive models)
- Human role: Human-designed; humans set goals and safety limits
- Verification: Peer-reviewed in Nuclear Fusion; hardware demonstrations
- Status: confirmed
- Sources: [PPPL: PACMAN AI framework makes key fusion decisions in milliseconds](https://www.pppl.gov/news/2026/pacman-ai-framework-controlling-fusion-systems-safely-makes-key-decisions-milliseconds) · [ScienceDaily: PACMAN AI framework for fusion](https://www.sciencedaily.com/releases/2026/09/260903064215.htm) · [Phys.org: PACMAN AI framework controls fusion systems safely](https://phys.org/news/2026-09-pacman-ai-framework-fusion-safely.html)
### 2026-02-13 — GPT-5.2 conjectures, and an OpenAI model proves, that 'single-minus' gluon tree amplitudes are nonzero
- Problem: Whether single-minus-helicity gluon tree amplitudes vanish
- Result: A closed-form formula and proof showing they are nonzero in half-collinear (2,2)-signature kinematics.
- AI system: GPT-5.2 Pro, OpenAI internal reasoning model
- Human role: Human-led with AI tools: physicists framed the question and verified
- Verification: Expert-checked by the author team; preprint
- Status: disputed
- Sources: [OpenAI: New result in theoretical physics](https://openai.com/index/new-result-theoretical-physics/) · [OpenAI: Extending single-minus amplitudes to gravitons](https://openai.com/index/extending-single-minus-amplitudes-to-gravitons/) · [Hugging Face blog: critical look at GPT and single-minus gluons](https://huggingface.co/blog/dlouapre/gpt-single-minus-gluons) · [The Quantum Insider: AI spots what physicists missed in gluon scattering](https://thequantuminsider.com/2026/02/13/ai-scientist-spots-what-physicists-missed-in-gluon-scattering/)
### 2025-12 — Physics Letters B paper built on a GPT-5 idea draws criticism that it tests the wrong thing
- Problem: Testing nonlinear modifications of quantum mechanics
- Result: Published criterion proposed by GPT-5; critics argue it does not test what it claims.
- AI system: GPT-5
- Human role: Human-led; AI supplied the main idea
- Verification: Peer-reviewed in Physics Letters B; disputed by experts
- Status: disputed
- Sources: [The Decoder: Physicist Steve Hsu publishes research built around a core idea generated by GPT-5](https://the-decoder.com/physicist-steve-hsu-publishes-research-built-around-a-core-idea-generated-by-gpt-5/) · [Oppenheim rebuttal (arXiv 2512.07809)](https://arxiv.org/abs/2512.07809) · [Peter Woit: Theoretical Physics Slop](https://www.math.columbia.edu/~woit/wordpress/?p=15362)
### 2025-10-16 — Google DeepMind partners with Commonwealth Fusion Systems to optimise and control the SPARC tokamak with AI
- Problem: Operating and controlling SPARC to reach net fusion energy (Q > 1)
- Result: Partnership and tooling announced; no fusion result yet.
- AI system: TORAX, reinforcement learning agents
- Human role: Human-led with AI tools
- Verification: Not applicable (partnership announcement)
- Status: pending
- Sources: [Google DeepMind: Bringing AI to the next generation of fusion energy](https://deepmind.google/blog/bringing-ai-to-the-next-generation-of-fusion-energy/) · [TORAX on GitHub](https://github.com/google-deepmind/torax)
### 2025-09-04 — DeepMind's Deep Loop Shaping cuts LIGO control noise 30–100×
- Problem: Low-frequency control noise limiting LIGO's sensitivity
- Result: Learned mirror-control policy reducing control noise by one to two orders of magnitude on real hardware.
- AI system: Deep Loop Shaping (RL)
- Human role: Human-designed; tested with LIGO engineers
- Verification: Peer-reviewed in Science; hardware demonstration
- Status: confirmed
- Sources: [Improving cosmological reach of a gravitational wave observatory using Deep Loop Shaping (Science)](https://www.science.org/doi/10.1126/science.adw1291) · [Caltech: Artificial intelligence helps boost LIGO](https://www.caltech.edu/about/news/artificial-intelligence-helps-boost-ligo)
### 2025-07-30 — Interpretable neural network discovers new non-reciprocal force laws in dusty plasma
- Problem: Inferring many-body force laws in dusty (complex) plasmas
- Result: Data-driven discovery of non-reciprocal force laws and corrections to standard charge and screening assumptions.
- AI system: physics-tailored neural network
- Human role: Human-led with AI tools: experiments and interpretation by physicists
- Verification: Peer-reviewed in PNAS; Cozzarelli Prize 2026
- Status: confirmed
- Sources: [ScienceDaily: AI just discovered new physics in the fourth state of matter](https://www.sciencedaily.com/releases/2026/04/260422044635.htm) · [Emory News: AI and dusty plasma](https://news.emory.edu/features/2025/07/esc_ai_dusty_plasma_30-07-2025/index.html) · [Emory: scientists receive Cozzarelli Prize](https://news.emory.edu/stories/2026/05/emory-scientists-receive-cozzarelli-prize-discovery-new-physics-dusty-plasma) · [arXiv 2310.05273 (preprint)](https://arxiv.org/abs/2310.05273)
### 2024-11-20 — AlphaQubit: neural decoder sets accuracy record for quantum error correction on Google's Sycamore
- Problem: Decoding surface-code error syndromes accurately
- Result: Most accurate decoder on real quantum hardware data at the time.
- AI system: AlphaQubit
- Human role: Human-designed model
- Verification: Peer-reviewed in Nature
- Status: confirmed
- Sources: [Learning high-accuracy error decoding for quantum processors (Nature)](https://www.nature.com/articles/s41586-024-08148-8) · [Google: AlphaQubit](https://blog.google/innovation-and-ai/models-and-research/google-deepmind/alphaqubit-quantum-error-correction/)
### 2024-02-21 — AI controller predicts and avoids tearing instabilities in the DIII-D fusion reactor
- Problem: Avoiding disruptive tearing-mode instabilities in high-performance tokamak plasmas
- Result: Real-time AI avoidance of tearing instabilities on a working tokamak.
- AI system: deep RL controller with learned dynamics model
- Human role: Human-designed; operated under human supervision
- Verification: Peer-reviewed in Nature; demonstrated on hardware
- Status: confirmed
- Sources: [Avoiding fusion plasma tearing instability with deep reinforcement learning (Nature)](https://www.nature.com/articles/s41586-024-07024-9) · [Princeton Engineering: Engineers use AI to wrangle fusion power](https://engineering.princeton.edu/news/2024/02/21/engineers-use-ai-wrangle-fusion-power-grid)
### 2022-02-16 — Deep reinforcement learning controls fusion plasma in the TCV tokamak
- Problem: Magnetic confinement and shaping of tokamak plasmas
- Result: First deep-RL controller to shape and sustain diverse plasma configurations on a real tokamak.
- AI system: deep reinforcement learning (MPO actor-critic)
- Human role: Humans built the simulator, specified targets and rewards, supervised experiments
- Verification: Peer-reviewed in Nature; demonstrated on hardware
- Status: confirmed
- Sources: [Magnetic control of tokamak plasmas through deep reinforcement learning (Nature)](https://www.nature.com/articles/s41586-021-04301-9) · [DeepMind: Accelerating fusion science through learned plasma control](https://deepmind.google/blog/accelerating-fusion-science-through-learned-plasma-control/)