Google unveils Gemini Omni, an any-to-any model that generates and conversationally edits video
Gemini Omni, announced at I/O on 19 May 2026, is Google's first "any-to-any" model family: Gemini Omni Flash takes text, images, audio and video in one prompt and outputs physics-aware video that can be edited turn-by-turn in plain language, with SynthID watermarks. It rolled out to paid Gemini/Flow users and free on YouTube Shorts; API access came 30 June and Omni 1.1 Flash on 27 Aug.
Key facts
- Announced 2026-05-19 at Google I/O; blog authored by Koray Kavukcuoglu
- Inputs: any mix of text, image, audio, video; first release (Omni Flash) outputs video only — image and audio output promised later
- Conversational editing keeps characters, lighting and continuity across turns; avatars with your own voice
- Rolled out to Google AI Plus/Pro/Ultra in Gemini app and Flow; free in YouTube Shorts Remix and YouTube Create (18+)
- SynthID watermark on every clip; speech-editing of real people restricted
- Developer API (gemini-omni-flash-preview) launched 2026-06-30; reported ~$0.10 per second of generated video
What happened
Instead of a standalone "Veo 4", Google introduced Gemini Omni, a generative model family that reasons across modalities rather than stitching separate models together. Gemini Omni Flash accepts a portrait, a location photo, a voice sample and a one-line brief in a single prompt and returns a single coherent shot; follow-up prompts edit the same scene. It shipped to consumers the same day and to Google Vids (Workspace) in July.
Why it matters
Omni folds Google's generative media stack (Veo, Nano Banana, Genie-style world knowledge) into the Gemini model line, and shifts video generation from one-shot prompting to iterative, conversational editing — a workflow closer to real production.
Changelog
- 2026-09-29: created
Videos (2)
Introducing Gemini Omni: Create Anything from Anything
Google · 2026-05-19 · officialDescription by Gemini, which watched the video:
Summary
This is an official promotional video produced by Google DeepMind showcasing the creative and generative capabilities of "Gemini Omni." Set to an upbeat instrumental track with no spoken voiceover, the video demonstrates multimodal video generation, real-time style transfers, scene modifications, and world building.
What is shown
- [00:00] Title card displaying "Gemini Omni" over natural spiral patterns (sunflower, chameleon tail, snail shell).
- [00:03] Text overlay "Create anything / From everything" displaying floating modality icons (audio, images, video, text prompts, 3D objects).
- [00:06] Video-to-video transformations of a man in front of a mirror: blowing fire, generating water ripples by touching glass, and transforming into a felt puppet, hand-drawn comic sketch, and voxel/block character.
- [00:14] Text "Look what you can do" across rapid scenes including a first-person whitewater kayak run, Martian landscape traversal in a space suit, a water slide, a desert stagecoach chase, and an animated pop-up sci-fi book.
- [00:19] Text "Build worlds" displaying material and structural swaps on a sculptural pavilion (illuminated patterns, yarn/knit texture, flower arches, foam bubbles) and liquid metal physics.
- [00:28] Motion-guided generation showing a drawn path that a 2D clownfish follows before leaping out of water into a realistic seascape.
- [00:31] Interface combining multimodal assets into a sci-fi scene, followed by contextual element editing: "Swap character" (astronaut replaced by a giant fish), "Swap detail" (space station ring replaced by flying origami cranes), "Swap style" (comic book line art), "Swap environment" (jungle planet canopy), and "Swap angle" (first-person helmet reflection).
- [00:42] Montage of diverse scenes including bio-architecture interiors, a lunar dome colony, skate video overlays ("POW!" comic effects), and UFOs descending over clouds.
- [00:48] Closing title cards displaying "Gemini Omni" over a black hole accretion disk and the "Google DeepMind" logo.
Claims & numbers
- None (the video contains no spoken claims, release dates, pricing, or quantitative benchmarks; on-screen copy consists solely of feature labels and promotional taglines).
Notable quotes
- [00:03] "Create anything / From everything" (on-screen text)
- [00:14] "Look what you can do" (on-screen text)
- [00:36] "Swap character / Swap detail / Swap style / Swap environment / Swap angle" (on-screen text)
Assessment
This is an official promotional teaser reel from Google DeepMind. The footage presents highly polished, cherry-picked visual outputs and conceptual editing capabilities rather than raw, unedited real-time interaction in an end-user UI.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.
Introducing Gemini Omni
Google for Developers · 2026-05-19 · officialDescription by Gemini, which watched the video:
Summary
In an episode of Google AI's Release Notes, host Logan Kilpatrick (Group Product Manager, AI Studio) is joined by Google DeepMind team members Nicole Brichtova, Dumitru Erhan, Gabe, and Shlomi Fruchter to introduce Gemini Omni (Gemini Omni Flash). The panel discusses and demonstrates the model's multimodal video generation and prompt-driven video editing capabilities, including character consistency, text rendering, audio synchronization, and safety features like SynthID watermarking.
What is shown
- Alphabet Rapid-Paced Sequence [02:07]: A generated stop-motion style clip cycling through the alphabet with handwritten letter slips and matching objects appearing in rapid temporal sequence (e.g., ball, egg, hat, key, quill, zipper).
- Video Editing / Subject Replacement [04:08 - 04:30]: A source video of a woman speaking is edited via prompt into an anthropomorphic wolf speaking with synchronized lip movements, expression nuance, and preserved original audio.
- Scene Transformation & Perspective Edits [08:18 - 09:14]: A violinist performing indoors is transported to an outdoor grass field based on reference images, subsequently modified to make her violin invisible, and then rendered from a reverse camera angle behind her shoulder.
- Physical & Stylistic Illusion Demos:
- A glass orb held in a hand reflecting an infinite checkered room [21:01].
- An open hand projecting a 3D topographic weather hologram displaying rendered text ("Tuesday, May 19 Mountain View, CA") [21:30, 21:39].
- A drawn marker circle on paper transitioning into an animated black hole sucking in tabletop items [28:47].
- An astronaut walking across terrain shifting through multiple artistic media (colored marker, sketch, 3D, retro comic) while preserving continuous motion [29:37].
- A claymation educational clip illustrating amino acid chains folding into alpha helices, beta sheets, and 3D proteins with voiceover and text titles [32:06].
- A pop-up papercraft storybook titled Sailor and the Sea with ambient lighting, animation, and voice narration [34:44].
- Personal Likeness & Voice Avatar Workflow [35:47, 36:07]: Video and audio generation reproducing Logan Kilpatrick's likeness and speech based on multi-angle reference photos and voice capture.
Claims & numbers
- Nicole Brichtova claims Gemini Omni brings "Nano Banana to video," combining multimodal inputs (image, video, audio, text) to generate video outputs, with more output modalities planned [00:56 - 01:23].
- Generation time for Gemini Omni clips is currently around 60 to 90+ seconds for a 10-second video output [10:04].
- Nicole states the model reliably follows instructions across 2 to 4 multi-turn edits [10:24].
- The avatar creation workflow supports uploading up to roughly 7 reference photos from multiple angles to improve 3D facial geometry understanding [27:00, 27:23].
- The model is available in the Gemini app for Ultra, Pro, and Plus users, in Google Flow for creative suites, and integrated into YouTube Shorts / YouTube Create for video remixing, with APIs coming soon [15:58, 16:21, 17:00, 17:10].
- All generated videos have SynthID invisible watermarks embedded directly into the video frames and include C2PA metadata, allowing detection via Google Chrome and the Gemini app [39:40 - 40:23].
Notable quotes
- Nicole Brichtova [00:56]: "One, is we're basically bringing Nano Banana to video. So we have a really great video generation model, but it especially shines at video editing."
- Shlomi Fruchter [02:41]: "The model has an ability to create very fast, potentially sequences... the control over the time and being able to tell a story is much better."
- Nicole Brichtova [15:57]: "It's available to Ultra and Pro and Plus users... So this is definitely a trade-off that we thought about with this model."
Assessment
This is an official Google DeepMind product showcase featuring panel discussion and pre-rendered demonstration reels. The showcased video generations illustrate strong temporal consistency, text rendering, and multimodal video editing, though the presenters acknowledge existing limitations including generation latency (60–90 seconds per 10-second clip), difficulty rendering large groups of people, and occasional over-editing when prompts are under-specified.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.
Related events
- Google I/O 2026: Gemini 3.5 Flash, Gemini Spark agent and Antigravity 2.0 ★★★★
- Gemini Omni Flash opens to developers via the Gemini API ★★
- Gemini Omni 1.1 Flash adds scene extension, frame interpolation and 4K upscaling ★★★
- Google DeepMind opens Project Genie, a Genie 3 world-model prototype, to AI Ultra subscribers ★★★
- Google launches Nano Banana 2 (Gemini 3.1 Flash Image) ★★★
Sources (6)
- officialIntroducing Gemini Omni (Google blog)
- officialGemini Omni Flash model card
- press9to5Google: Gemini Omni, the 'create anything' model
- pressTechCrunch: Gemini Omni turns images, audio and text into video
- videoIntroducing Gemini Omni: Create Anything from Anything (video)
- officialGemini Omni Flash now in Google Vids (Workspace blog)
id: 2026-05-19-gemini-omni · updated 2026-09-29 · open in the interactive timeline