Gemini Robotics ER 2
Vision-language model for robotics (outputs text/JSON, not motor commands). 131,072 input / 65,536 output tokens. Standard id supports caching, code execution, computer use, file search, function calling, Search and Maps grounding, structured outputs and thinking. Replaces gemini-robotics-er-1.6-preview (shut down 2026-08-31). No GA id yet. Knowledge cutoff not stated.
- Context window
- 131,072 tokens
- Max output
- 65,536 tokens
- Input
- text, image, video, audio
- Output
- text
- License
- proprietary
- Pricing
- input: $1 · output: $5 (per 1M tokens (text/image/video/audio input); introductory rate through 2026-12-31, rising to $2.00 in / $10.00 out from 2027-01-01; Batch API half price) source
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Gemini API | gemini-robotics-er-2-preview | https://generativelanguage.googleapis.com/v1beta/models/gemini-robotics-er-2-preview:generateContent | docs |
| Gemini API (Live API, streaming) | gemini-robotics-er-2-streaming-preview | — | docs |
| Google AI Studio | — | aistudio.google.com | — |
| Gemini Enterprise Agent Platform (Google Cloud, private preview) | — | — | docs |
| Sample code (GitHub) | — | github.com/google-gemini/robotics-samples | — |
Notable capabilities (4)
- Embodied reasoning "robot brain" in a public API: Spatial reasoning (points, boxes, trajectories), multi-step task planning, tool/function calling and code execution to orchestrate a robot's VLA or controller; publicly callable, unlike the VLA models. source
- Continuous video monitoring and task-progress tracking: Watches video feeds to track progress and adapt; Google reports 91.3% moment-finding accuracy (0.96 s mean absolute distance) at ~4x the speed of the previous generation and 57.4% progress classification. source
- Low-latency streaming via Live API: Separate gemini-robotics-er-2-streaming-preview id supports bidirectional audio/video streaming with function calling and thinking (no caching, code execution or structured output). source
- Multi-robot collaboration: Coordinates heterogeneous robots (e.g. wheeled rovers and humanoids, Boston Dynamics Spot demo) to communicate and hand off tasks. source
The hosted, publicly callable half of the Gemini Robotics 2 stack: a high-level planner that points at objects, plans multi-step tasks and calls a robot's own VLA/skills as tools.
Sources: model page, pricing, Google blog, model card.
Timeline entry
- Google DeepMind launches Gemini Robotics 2 family with whole-body humanoid control ★★★★
On 30 July 2026 Google DeepMind released Gemini Robotics 2 — a VLA model for whole-body humanoid control and dexterous manipulation, the Gemini Robotics ER 2 embodied-reasoning "brain" (public in the Gemini API), and a lightweight On-Device 2 model that adapts to new robot bodies in hours…
Other Google DeepMind models
Gemini 3.8 Flash TTS · Gemini 3.8 Live · Gemini 3.8 Flash · Gemini 3.5 Transcribe (and Transcribe Live) · Lyria 3.5 · Gemini 3.5 Flash-Lite · Nano Banana 2 Lite (Gemini 3.1 Flash-Lite Image) · Gemini Omni Flash (Omni 1.1 Flash) · Gemini Embedding 2 · Gemma 4 · Nano Banana 2 (Gemini 3.1 Flash Image) · Nano Banana Pro (Gemini 3 Pro Image) · Gemini Robotics 2 · Gemini Robotics On-Device 2 · Gemini 3.5 Live Translate · Gemini 3.1 Pro · Veo 3.1 · Genie 3 · Lyria RealTime · Gemini 3.7 Flash · Gemini 3.6 Flash · Gemini 3.5 Flash · Gemini 3.1 Flash TTS (preview) · Gemini 3.1 Flash Live (preview) · Lyria 3 (Clip / Pro) · Lyria 2 · Gemini 2.5 Flash Native Audio (Live, preview) · Gemini 2.5 Flash-Lite · Gemini 2.5 Flash · Gemini 2.5 Pro · Gemini 2.5 Flash TTS / Pro TTS · Gemini 3.1 Flash-Lite · Nano Banana (Gemini 2.5 Flash Image) · Gemini Robotics-ER 1.5 / 1.6 · Imagen 4