Introducing EmbeddingGemma 2: An open model for natively multimodal embeddings
Google for Developers · 2026-10-06 · official · 120,234 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Google product managers Sebastian Russo and Ivan Llanos, alongside research lead Sahil Dua, introduce EmbeddingGemma 2, an open, lightweight model family for natively multimodal embeddings. The video details its unified embedding architecture, parameter sizes, benchmark performance, and on-device deployment capabilities across mobile and laptop environments.
What is shown
- Unified Embedding Visualization [00:16]: Conceptual animation showing text, images, video, and audio mapped into a single high-dimensional embedding space (e.g., the word "cat", picture of a cat, and cat meow sound clustered together).
- Benchmark Visualizations [00:45]: Charts comparing EmbeddingGemma 2 against alternatives on the MIEB (lite) and MAEB benchmarks across model sizes.
- Architecture & Modularity Breakdown [00:54]: Graphic showing the modular configurations (270M text, 440M text + vision, 570M text + audio, and 740M text + vision + audio), Matryoshka dimension scaling (128, 256, 512, 768 dimensions), and an 8K token context window.
- Google AI Edge Gallery App Demo [01:26]: On-device search on a smartphone using natural language queries ("t shirt") to instantly locate matching photos, followed by Video Moment Finder locating the timestamp where candles are blown out from the query "Blowing out candles" without transcription [01:39].
- AI Edge Foresight RAG Demo [02:16]: On-laptop meeting assistant running locally; EmbeddingGemma 2 embeds live spoken audio, detects a spoken query ("How does foresight actually leverage embedding Gemma 2?"), retrieves a local architecture diagram, and pairs with Gemma 4 to answer in real time.
- Fine-Tuning Use Cases [02:50]: Visuals highlighting domain-specific fine-tuning for legal contracts, medical imaging, and technical product catalogs.
Claims & numbers
- Model Configurations: Sahil Dua states the model comes in modular form factors: 270M (Text), 440M (Text + Vision), 570M (Text + Audio), and up to 740M parameters (Text + Vision + Audio) [00:54].
- Dimensions & Truncation: Dua states the model natively outputs 768-dimensional vectors, but Matryoshka representation learning allows truncation down to 512, 256, or 128 dimensions [01:02].
- Context Length: Dua states the model features an 8,000 (8K) token context window [01:12].
- Performance: Dua claims EmbeddingGemma 2 sets a new standard for models under 1 billion parameters on MIEB (lite) and outperforms much larger models on several MAEB benchmarks [00:45].
- Privacy & Local Execution: Llanos and Russo claim that search, retrieval, and RAG pipelines execute entirely on-device without intermediate transcriptions, cloud API calls, or sensitive data leaving the hardware [01:48, 02:08].
- Availability: Llanos states model weights are available on Hugging Face and notebooks are in the Gemma Cookbook [03:00].
Notable quotes
- [00:00] Sebastian Russo: "Today we're super excited to introduce EmbeddingGemma 2, a lightweight, open model for natively multimodal embeddings."
- [00:23] Ivan Llanos: "EmbeddingGemma 2 changes that by mapping text, images, video, and audio, or any combination of them, into a single, unified embedding space."
- [01:48] Ivan Llanos: "This all happens directly on my phone without any intermediate transcription, and no external API calls are required."
Assessment
This is an official Google launch and technical overview video showcasing EmbeddingGemma 2's capabilities, specifications, and on-device use cases. The mobile and laptop demonstrations reflect functional prototype applications (AI Edge Gallery and AI Edge Foresight), though the video editing cuts between slides and pre-recorded UI captures rather than showing full unedited continuous interactions.
Described by gemini-3.8-flash on 2026-10-07 from the video's audio and frames.