Introducing Gemini Omni
Google for Developers · 2026-05-19 · official · 18,260 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
In an episode of Google AI's Release Notes, host Logan Kilpatrick (Group Product Manager, AI Studio) is joined by Google DeepMind team members Nicole Brichtova, Dumitru Erhan, Gabe, and Shlomi Fruchter to introduce Gemini Omni (Gemini Omni Flash). The panel discusses and demonstrates the model's multimodal video generation and prompt-driven video editing capabilities, including character consistency, text rendering, audio synchronization, and safety features like SynthID watermarking.
What is shown
- Alphabet Rapid-Paced Sequence [02:07]: A generated stop-motion style clip cycling through the alphabet with handwritten letter slips and matching objects appearing in rapid temporal sequence (e.g., ball, egg, hat, key, quill, zipper).
- Video Editing / Subject Replacement [04:08 - 04:30]: A source video of a woman speaking is edited via prompt into an anthropomorphic wolf speaking with synchronized lip movements, expression nuance, and preserved original audio.
- Scene Transformation & Perspective Edits [08:18 - 09:14]: A violinist performing indoors is transported to an outdoor grass field based on reference images, subsequently modified to make her violin invisible, and then rendered from a reverse camera angle behind her shoulder.
- Physical & Stylistic Illusion Demos:
- A glass orb held in a hand reflecting an infinite checkered room [21:01].
- An open hand projecting a 3D topographic weather hologram displaying rendered text ("Tuesday, May 19 Mountain View, CA") [21:30, 21:39].
- A drawn marker circle on paper transitioning into an animated black hole sucking in tabletop items [28:47].
- An astronaut walking across terrain shifting through multiple artistic media (colored marker, sketch, 3D, retro comic) while preserving continuous motion [29:37].
- A claymation educational clip illustrating amino acid chains folding into alpha helices, beta sheets, and 3D proteins with voiceover and text titles [32:06].
- A pop-up papercraft storybook titled Sailor and the Sea with ambient lighting, animation, and voice narration [34:44].
- Personal Likeness & Voice Avatar Workflow [35:47, 36:07]: Video and audio generation reproducing Logan Kilpatrick's likeness and speech based on multi-angle reference photos and voice capture.
Claims & numbers
- Nicole Brichtova claims Gemini Omni brings "Nano Banana to video," combining multimodal inputs (image, video, audio, text) to generate video outputs, with more output modalities planned [00:56 - 01:23].
- Generation time for Gemini Omni clips is currently around 60 to 90+ seconds for a 10-second video output [10:04].
- Nicole states the model reliably follows instructions across 2 to 4 multi-turn edits [10:24].
- The avatar creation workflow supports uploading up to roughly 7 reference photos from multiple angles to improve 3D facial geometry understanding [27:00, 27:23].
- The model is available in the Gemini app for Ultra, Pro, and Plus users, in Google Flow for creative suites, and integrated into YouTube Shorts / YouTube Create for video remixing, with APIs coming soon [15:58, 16:21, 17:00, 17:10].
- All generated videos have SynthID invisible watermarks embedded directly into the video frames and include C2PA metadata, allowing detection via Google Chrome and the Gemini app [39:40 - 40:23].
Notable quotes
- Nicole Brichtova [00:56]: "One, is we're basically bringing Nano Banana to video. So we have a really great video generation model, but it especially shines at video editing."
- Shlomi Fruchter [02:41]: "The model has an ability to create very fast, potentially sequences... the control over the time and being able to tell a story is much better."
- Nicole Brichtova [15:57]: "It's available to Ultra and Pro and Plus users... So this is definitely a trade-off that we thought about with this model."
Assessment
This is an official Google DeepMind product showcase featuring panel discussion and pre-rendered demonstration reels. The showcased video generations illustrate strong temporal consistency, text rendering, and multimodal video editing, though the presenters acknowledge existing limitations including generation latency (60–90 seconds per 10-second clip), difficulty rendering large groups of people, and occasional over-editing when prompts are under-specified.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.