Google Just Turned Street View Into a Video Game
Bilawal Sidhu · 2026-05-19 · ai-made · 245,785 views
Made by AI
Model: Genie 3 · Series: World-model footage
Evidence: Description: Genie 3 with Maps Imagery Grounding can 'generate interactive 3D worlds conditioned to any of the 280 billion Street View images'; early-access demo.
Human role: Bilawal Sidhu (former Google Maps PM) picked the places and styles and played them.
Pipeline: Google Maps location → Genie 3 world model (Street View grounded) → real-time playable world, screen-recorded
Lore: world-model-walk
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
In this video, creator and former Google Maps product lead Bilawal Sidhu reviews Google DeepMind’s Project Genie (Genie 3) integration with Google Maps Street View imagery, announced around Google I/O. He demonstrates how interactive real-time world-generation models can turn 360-degree Street View panoramas into playable, editable 3D-like simulation environments.
What is shown
- [00:00 - 00:44] Introduction to grounding Genie 3 experiences using Google Street View panoramic imagery, showing early demo clips (raccoon on a scooter, Formula 1 car, runner in Austin).
- [00:45 - 00:52] The Project Genie interface showing prompt fields for Environment ("Choose a location from Google Maps") and Character, with a third-person camera toggle.
- [00:53 - 01:29] Driving simulation of a Google Maps-themed Formula 1 car navigating the Las Vegas Strip, complete with an AI-generated speedometer HUD, race checkpoints, and Parisian landmarks.
- [01:30 - 02:02] Third-person simulation of a raccoon and a fox riding scooters around and through the Palace of Fine Arts in San Francisco.
- [02:03 - 02:23] Simulation featuring Google Maps mascot Pegman running past the Ferry Building in San Francisco.
- [02:24 - 03:12] An avatar running along the Ann and Roy Butler Hike-and-Bike Trail over Lady Bird Lake in Austin, Texas, jumping over a railing into the water, and switching to a boat simulation under railway bridges.
- [03:13 - 03:20] Indoor walkthrough of the White House generated from indoor Street View "special collects."
- [03:21 - 03:49] Conceptual transformations, including underwater scuba diving beneath the Golden Gate Bridge, snowstorms on city streets, and historical black-and-white aerial imagery.
- [03:50 - 04:15] Discussion of world models illustrated by a Spider-Man pointing meme representing competing approaches (JEPA, LLM, SLAM, Video-Gen, 3DGS, Google Maps).
- [04:16 - 05:40] Breakdown of retrieval-augmented generation (RAG) for world models using the "Seoul World Model" academic paper as an architectural comparison.
- [06:31 - 06:45] A TechCrunch quote from Jack Parker-Holder noting real-time models lag offline video models by roughly 6 to 12 months in quality.
Claims & numbers
- The presenter states that Genie 3 is Google's real-time interactive world model that autoregressively generates the next video frame based on user controls and inputs.
- The presenter notes that the current version of Project Genie relies only on Street View panoramic photography rather than aerial imagery.
- The presenter quotes Jack Parker-Holder (from a TechCrunch article) stating that this kind of interactive world model is "maybe six to 12 months behind video in terms of the accuracy and quality."
Notable quotes
- [00:19] "What that means is you can reference actual Street View photography of a physical area and use that as a basis for your generation."
- [01:13] "And this is particularly cool because this is just referencing the panoramic imagery. They're not even feeding in the aerial imagery into it yet."
- [06:34] "'I think for this kind of model, it's maybe six to 12 months behind video in terms of the accuracy and quality, so I think it's something we will solve,' Parker-Holder said."
Assessment
This is a creator review and demonstration video examining early access to Google DeepMind's Project Genie Street View integration. The interactive gameplay sequences are actual prototype screen recordings from Genie 3, highlighting both impressive dynamic generation and noticeable visual hallucination artifacts when deviating far from original camera angles.
Lyrics & themes
This video is spoken commentary and demonstration rather than a song. The narration revolves around turning physical mapping data into real-time interactive virtual simulations:
- Real-world holodeck: "How do you take the complexity of reality and put it inside a simulation so you can do anything inside it?" [00:03]
- Interactive generation: "This model is autoregressively predicting the next frame... it can just generate everything on the fly for you." [01:50]
- World simulation editing: "So kind of by bringing reality into latent space, you can now edit it and do things that would have been otherwise very hard or tedious to do in traditional tools." [05:03]
- The future of game engines: "Is this what you imagine GTA 7 is actually going to look like?" [07:33]
Lore & references
- Pegman: The yellow human-shaped icon from Google Maps, animated here as a playable 3D character exploring San Francisco.
- World Models Meme: A classic multi-Spider-Man meme highlighting the rivalry between different paradigms for digital reality representation: Meta's JEPA, LLMs, robotics SLAM, generative video models, 3D Gaussian Splatting (3DGS), and geospatial datasets like Google Maps.
- Seoul World Model (SWM): Reference to a research paper on retrieval-augmented generation (RAG) conditioning video diffusion models on city-scale Street View databases.
- GTA 7: A running gaming culture reference speculating that neural world models will eventually replace traditional polygon-based game engines in future open-world titles.
Visual style & craft
The video blends standard creator video essay production—a lighted webcam talking-head shot and screen recordings of web articles and X (Twitter) threads—with direct gameplay captures of Google’s Genie 3 neural world simulator. The generated simulations exhibit characteristic neural video artifacts, including edge warping, object morphing when pivoting cameras, and dreamlike background hallucinations, contrasting with the static, crisp 2D UI overlays and web interfaces.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.