Post-Cutoff

Review

OpenAI’s 722 AI Math Papers, and Why Terence Tao Is Uneasy

Simply AIYouTube90,674 views as of 10 October 2026

Watch on YouTubePlay loads YouTube’s player from youtube-nocookie.com.

Description

Description written by Gemini from the videoGemini 3.8 Flash, 10 October 2026

Summary

The video by Simply AI examines OpenAI’s massive repository release containing 722 AI-generated math manuscripts organized into 372 result families. It discusses the mixed reaction from the mathematical community—including cautionary commentary from Terence Tao—and examines how automated proof verification with Lean works in practice, detailing a personal attempt to build and verify a Lean formalization on a Mac.

What is shown

  • [00:00] Overview of OpenAI’s GitHub repository openai/math, displaying the README, paper categories, and mentions of unreleased internal OpenAI models.
  • [00:15] Visual breakdown of Lean verification coverage across the collection (330 of 722 papers verified via Lean; 392 without Lean formalization).
  • [00:30] Social media reaction from mathematician Terence Tao warning of an “unsustainable fashion” of harvesting open problems.
  • [00:54] Terminal demonstration building one of the Lean formalizations (lake build OAI.NumberTheory.PiExponent.Main).
  • [00:57] Mathematical explanation and animated visualization of the Flint-Hills series and irrationality exponent of $\pi$.
  • [02:43] Excerpts from OpenAI’s 42-page summarized chain of thought document detailing the model’s exploratory steps and dead ends.
  • [03:27] Code walkthrough in Lean 4 (PIExponent.lean), explaining how the top-level theorem statement uses the sorry keyword while delegating to hundreds of supporting formalization files.
  • [05:34] Benchmark timeline and performance stats from running the Lean build locally on a Mac, checking 630 of 869 files over 84 minutes with zero errors before halting.
  • [06:24] Analysis of the “five receipts” recommended by the Institute for Advanced Study advisory group compared against what OpenAI actually released.

Claims & numbers

  • The presenter states OpenAI released 722 manuscripts organized into 372 families on GitHub, generated by an unreleased internal OpenAI model.
  • The presenter states 330 of the 722 papers have Lean verification checks, while 392 have no formal verification check.
  • The presenter cites Alex Kontorovich stating that if a human had achieved the quasi-Riemann result, “it would be an instant Fields Medal.”
  • The presenter states that for the average result, the model required roughly three hours of ChatGPT Pro thinking compute, having been posed approximately 4,000 problems (yielding about 370 result families, a solve rate under 1 in 10).
  • The presenter claims the best mathematical bound for the irrationality exponent of $\pi$ prior to this work was $7.1$ (from 2020), which the AI claims to bring down to exactly $2.0$.
  • The presenter states the complete Lean proof for the $\pi$ exponent theorem consists of 869 files and approximately 95,000 lines of formal code, which was also written by an AI agent.
  • The presenter ran lake build on a Mac: 630 of 869 files completed verification in 84 minutes without errors before the run was stopped, with the slowest individual files taking 10 to 17 minutes each.
  • The presenter notes that of the five transparency “receipts” requested by the IAS-hosted advisory group, OpenAI provided reasoning summaries for only 10 results, gave only an average compute time (no exact dollar figure), and did not release the prompts or the model’s name.

Notable quotes

  • [00:35] Terence Tao: “solutions to open problems are now being harvested at large scale in an unsustainable fashion”
  • [00:43] Alex Kontorovich: “If a human did this, it would be an instant Fields Medal.”
  • [06:19] Andrew Sutherland: “We should ask for receipts.”

Assessment

This is an independent analysis and hands-on technical review of OpenAI’s open-source math proof release. The presenter performs a genuine local reproduction of Lean compilation while providing a grounded, critical evaluation that clearly distinguishes between what Lean actually verifies (internal syntactic consistency) and what requires human review (the fidelity of the formalization to the claimed mathematical statement).

Described by gemini-3.8-flash on 2026-10-10 from the video’s audio and frames.

Related

  1. Science & math 98 days after the cutoff

    OpenAI releases 722 AI-written math manuscripts claiming hundreds of open problems