EmbeddingGemma 2
Announced 2026-10-06 on the Google blog. Apache 2.0 (Gemma 4 license), not gated on Hugging Face. Runs on phones and laptops; uses task-instruction prefixes (e.g. 'task: search result | query: ...') for text inputs. sentence-transformers id: google/embeddinggemma-2. No hosted Gemini API endpoint is listed; for hosted multimodal embeddings Google offers gemini-embedding-2.
- Context window
- 8,192 tokens
- Input
- text, image, video, audio
- Output
- text
- License
- apache-2.0
- Verified
- 2026-10-07
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Hugging Face | google/embeddinggemma-2 | huggingface.co/google/embeddinggemma-2 | — |
| Google docs | — | — | docs |
Notable capabilities (3)
- Four modalities in one small open embedding space: Maps text (including code), images, video and audio into one 768-dimensional space with 740M parameters in total: a 270M text model (130M backbone + 140M embedder) plus optional 170M vision and 300M audio encoders that can be loaded separately. source
- Matryoshka truncation: Embeddings can be cut to 512, 256 or 128 dimensions (up to 6x less storage); quality loss is small down to 256d. source
- Big code-retrieval gain over v1: MTEB code v1 78.68 vs 68.76 for EmbeddingGemma 1; MTEB multilingual v2 61.36 vs 61.15; MMEB v2 overall 59.01 (Google's numbers). source
Small open on-device embedding model for multimodal search, RAG, classification and clustering.
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("google/embeddinggemma-2")
q = model.encode("What causes the northern lights?", prompt_name="SearchQuery")
d = model.encode("The northern lights are caused by charged particles from the sun.", prompt_name="Document")
print(model.similarity(q, d))
Sources: Hugging Face model card, launch blog, docs.
Timeline entry
- Google DeepMind releases EmbeddingGemma 2: a 740M-parameter open (Apache 2.0) embedding model for text, code, images, video and audio ★★★
On 6 Oct 2026 Google DeepMind released EmbeddingGemma 2, an open Apache 2.0 embedding model that maps text (including code), images, video and audio into one 768-dimensional space with 740M parameters in total, small enough for phones and laptops. It reached 271 points on Hacker News.
Other Google DeepMind models
Nano Banana 2.1 (Gemini Nano Banana 2.1) · Gemini 3.8 Flash TTS · Gemini 3.8 Live · Gemini 3.8 Flash · Gemini 3.5 Transcribe (and Transcribe Live) · Lyria 3.5 · Gemini 3.5 Flash-Lite · Nano Banana 2 Lite (Gemini 3.1 Flash-Lite Image) · Gemini Omni Flash (Omni 1.1 Flash) · Gemini Embedding 2 · Gemma 4 · Nano Banana Pro (Gemini 3 Pro Image) · Gemini 4 Argon · Gemini Robotics 2 · Gemini Robotics ER 2 · Gemini Robotics On-Device 2 · Gemini 3.5 Live Translate · Gemini 3.1 Pro · Veo 3.1 · Genie 3 · Lyria RealTime · Gemini 3.7 Flash · Gemini 3.6 Flash · Gemini 3.5 Flash · Gemini 3.1 Flash TTS (preview) · Gemini 3.1 Flash Live (preview) · Lyria 3 (Clip / Pro) · Lyria 2 · Gemini 2.5 Flash Native Audio (Live, preview) · Gemini 2.5 Flash-Lite · Gemini 2.5 Flash · Gemini 2.5 Pro · Gemini 2.5 Flash TTS / Pro TTS · Gemini 3.1 Flash-Lite · Nano Banana 2 (Gemini 3.1 Flash Image) · Nano Banana (Gemini 2.5 Flash Image) · Gemini Robotics-ER 1.5 / 1.6 · Imagen 4