Create your own voices with Gemini 3.8 text-to-speech
Google DeepMind · 2026-09-23 · official · 140,022 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
This official Google DeepMind promotional video demonstrates the voice-creation capabilities of Gemini's text-to-speech technology (Gemini Audio). A presenter walks through prompting a custom voice persona, browsing a voice library, remixing voice attributes, building a multi-speaker dialogue scene, and replicating his own voice with a verification check.
What is shown
- Voice Design from prompt [00:04–00:18]: The presenter enters the prompt "Woman in her 30's, pumped, emphatic with Brooklyn accent", generating several audition options and selecting one ("Gemini Audio Promo Narrator").
- Browse Voice Library [00:19–00:26]: Navigating a library of preset personas, including "The Mad Scientist", "The Meditation Guide", "The Hype Announcer", "The Ad Voiceover", "The Late-Night DJ", and "The Cheeky Sidekick".
- Voice Remixing [00:27–00:33]: Modifying a selected voice with the text prompt "Make this voice a little bit deeper", generating deeper variations (labeled with an on-screen note: "Voice remix coming soon").
- Multi-speaker dialogue editor [00:34–00:43]: Arranging two generated voices into a multi-turn conversation script with inline direction tags (e.g.,
[uh huh]). - Voice Replication & Verification [00:45–00:54]: Recording ~20 seconds of speech reading a sample script ("The lighthouse keeper watched the silver horizon..."), followed by a spoken consent verification phrase ("I am the owner of this voice and consent to Google using this voice to create a synthetic voice model"), generating a synthetic clone ("Luke's voice- 1").
- End screen [00:55–01:00]: The tagline "Give your words a voice" alongside the Gemini Audio logo.
Claims & numbers
- The narrator claims the library includes "over a thousand ready-to-go voices" [00:24].
- The Voice Replication interface states that "twenty seconds of natural speech works best" to create a clone [00:46].
- The voiceover claims synthetic voice creation is "safely protected by a quick verbal identity check" [00:52].
Notable quotes
- "Transforms simple prompts into consistent, expressive voice outputs on demand." [00:13]
- "Or choose from over a thousand ready-to-go voices." [00:24]
- "Give your words a voice with Gemini Audio." [00:55]
Assessment
This is a polished official promotional demonstration by Google DeepMind. Small-print disclaimers disclose that on-screen sequences are shortened and simulated for illustrative purposes, and the voice remix feature is marked as "coming soon".
Described by gemini-3.8-flash on 2026-10-06 from the video's audio and frames.