Post-Cutoff

Review

Did Google just kickstart the intelligence explosion?

FireshipYouTube1,975,806 views as of 8 October 2026

Watch on YouTubePlay loads YouTube’s player from youtube-nocookie.com.

Why it is here

Fireship (Sept 17): ‘Google DeepMind just published Dream-RSI, a technique that turns an AI’s old discovery logs into a simulator so it can test thousands of exploration strategies.’ ~1.98M views. Dream-RSI has no timeline entry yet (added to leads on 2026-10-08); unverified beyond this description. Length 4:58.

Description

Description written by Gemini from the videoGemini 3.8 Flash, 8 October 2026

Summary

In this episode of The Code Report, host Jeff Delaney examines recent developments in recursive self-improvement (RSI) for AI, focusing on research papers from ByteDance/Tsinghua University and Google DeepMind/University of Maryland (“Dream-RSI”). He breaks down how Dream-RSI uses past discovery logs to simulate and optimize search/exploration policies without modifying the underlying model weights, questioning whether this represents genuine recursive self-improvement or advanced search optimization. The video also features a sponsored demonstration of Blacksmith’s GitHub Actions runners and its new cloud coding agent, Codesmith.

What is shown

  • [00:00 - 00:24] Historical context on Irving John Good’s 1965 paper (“Speculations Concerning the First Ultraintelligent Machine”) and the concept of “seed AI” and recursive self-improvement (RSI).
  • [00:25 - 00:41] Overview of the paper “The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement” by 33 researchers from ByteDance, Tsinghua University, and other labs, outlining a 5-level roadmap (L1 Execution to L5 Meta-Improvement).
  • [00:42 - 01:10] Introduction to Google DeepMind and University of Maryland’s paper “Dream-RSI: Recursive Self-Improvement through Evolving Worlds”, and community debate on whether policy search qualifies as true RSI.
  • [01:14 - 01:37] List of summer 2026 mathematical conjectures solved or addressed by frontier models (e.g., GPT-5.6 Sol, Claude Fable 5, GPT-6 Astra).
  • [01:38 - 02:08] Explanation of evolutionary coding loops (like AlphaEvolve) and the concept of “Exploration Policy” deciding search paths (exploit vs. explore).
  • [02:09 - 02:37] Architectural explanation and demo animation of Dream-RSI constructing a replay simulator from cached historical traces to “dream” up alternative exploration policies at zero execution cost.
  • [02:40 - 03:15] Benchmark results from the Dream-RSI paper, including algorithm design and Lasso solver optimization tasks, along with prompt engineering excerpts used in policy evolution.
  • [03:16 - 04:06] Critical analysis of whether Dream-RSI fits Good’s definition of RSI, concluding that the underlying model (Gemini) remains unchanged and merely refines its search trajectory within its existing capability bounds.
  • [04:07 - 04:52] Sponsored walkthrough of Blacksmith CI runners and Codesmith agent resolving a pull request across multiple repositories via Slack and GitHub.

Claims & numbers

  • The presenter states that I.J. Good wrote in 1965 that the first ultraintelligent machine would be “the last invention man need ever make.” [00:09]
  • Last week, 33 researchers from ByteDance, Tsinghua University, and other labs published “The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement” featuring a 5-stage roadmap. [00:25]
  • The presenter notes that Google DeepMind and University of Maryland published “Dream-RSI: Recursive Self-Improvement through Evolving Worlds”. [00:42]
  • In benchmark tests against eight algorithm design and math problems, the presenter notes Dream-RSI wrote a Lasso solver that outperforms Python’s standard scikit-learn library in about 300 attempts, compared to 550 attempts for a fixed static policy and 51,200 generations for the previous record holder (SimpleTES). [02:51]
  • The presenter claims Blacksmith’s GitHub Actions runners run twice as fast while costing 75% less. [04:12]

Notable quotes

  • “Let an ultraintelligent machine be defined as a machine that can far surpass all the intellectual activities of any man however clever... Thus the first ultraintelligent machine is the last invention that man need ever make...” [00:09] (quoting I.J. Good)
  • “Only the exploration-policy code changes; the underlying models, evaluator, and execution interfaces remain fixed.” [00:56] (quoting the Dream-RSI paper)
  • “The model that writes each new exploration policy is still the same Gemini, so it can never find a solution that it wasn’t already capable of writing; it just finds them faster and with fewer wasted attempts.” [03:25]

Assessment

This is an analytical tech commentary and review video by Fireship, summarizing two recent academic papers on recursive self-improvement and putting them in context with recent AI-assisted math breakthroughs. While the presenter relies on humor and memes, the technical explanation of Dream-RSI’s replay simulator and the distinction between weight self-improvement and policy-search optimization are accurately presented based on the published research, followed by a real promotional demo of Blacksmith and Codesmith.

Described by gemini-3.8-flash on 2026-10-08 from the video’s audio and frames.