Introducing GEN-1.5, a one-shot learner
Generalist · 2026-08-19 · official · 195,546 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
This official launch video from Generalist AI introduces GEN-1.5, a robot foundation model designed as a "one-shot learner" capable of immediate physical in-context learning. Through a narrated overview and laboratory footage, the company showcases dual-arm manipulator robots learning new manipulation tasks within seconds from short demonstrations, simulation data, and direct human hand gestures without task-specific retraining.
What is shown
- In-Context and Few-Shot Learning Demos [00:14–00:40]: Bimanual robotic arms equipped with customized multi-finger grippers unzipping pouches, stacking cups, opening jars, folding paper, and transferring behaviors learned from simulator prompts to physical hardware.
- Few-Shot Task Performance Chart [00:41–00:47]: Benchmark results showing task success rates when fine-tuned on 10 gradient steps (~5 minutes of data).
- Physical Prompting Architecture [01:01–01:16]: Conceptual schematic illustrating how prompt frames and live sensor input frames are passed into the model weights to generate robot trajectories without gradient updates.
- In-Context vs. Few-Shot Comparison Chart [01:31–01:50]: Benchmark comparisons showing zero-gradient in-context learning (3–12 seconds of prompt demos) achieving 37%–78% success across 10 distinct manipulation tasks, compared to 10-step fine-tuning.
- Novel Tool Use Improvisation [01:52–02:25]: A robot using an actual banana to sweep a cube into a bowl [02:01], using a dustpan and opposite arm cooperatively to scoop and dump objects [02:11], and switching tools ambidextrously.
- Improvisational Problem-Solving [02:26–02:57]: The robot dislodging a Lego brick stuck to its gripper with its other hand [02:34], removing a sheet of paper obstructing a bowl before dropping an object in [02:38], and adapting single-hand unscrewing techniques to two hands across various bottle and cup types [02:47].
- Human-to-Robot In-Context Learning [02:58–03:24]: An engineer demonstrates cup stacking with bare hands directly in front of the robot, which immediately replicates the stacking sequence on its own cups.
Claims & numbers
- The narrator claims GEN-1.5 can learn and generalize new tasks in seconds using physical in-context prompting with zero training/gradient updates on the target task.
- In few-shot mode (10 gradient steps / 5 minutes of data), reported success rates include:
- Sweep Trash With Brush: 99%
- Twist Lid Off Glass Jar: 94.5%
- Remove Vacuum Pad: 96%
- Unzip Pencil Pouch: 86%
- Retrieve Money From Wallet: 83.3%
- Open Book Cover: 82.7%
- Flip Phone Upside Down: 81%
- Stack Two Small Cups: 75%
- Brush Cube Into Bowl: 71.2%
- Fold and Crease Paper: 69.3%
- In zero-shot/in-context mode (3–12 seconds of demonstration), reported success rates include:
- Flip Phone Upside Down: 78%
- Stack Two Small Cups: 67%
- Remove Vacuum Pad: 64%
- Retrieve Money From Wallet: 60.7%
- Brush Cube Into Bowl: 60.8%
- Twist Lid Off Glass Jar: 60%
- Unzip Pencil Pouch: 55.5%
- Open Book Cover: 54.7%
- Fold and Crease Paper: 50%
- Sweep Trash With Brush: 37.3%
- On a held-out validation task, 0-step in-context learning scored 67%, 1 step scored 66.5%, 5 steps scored 58%, and 10 steps reached 75%.
Notable quotes
- "Our new model, GEN-1.5, is an immediate learning generalist. It's a one-shot learner." [00:15]
- "The fastest way it learns is with zero training on a new task, and just a few seconds of demonstration data put into the model's context." [01:01]
- "We're also starting to see human-to-robot in-context learning emerge, where a person can just show the robot what to do with their own human hands, and the robot mimics it on the spot with its hands." [02:59]
Assessment
This is an official demonstration and announcement video combining real lab footage, system diagrams, and evaluation charts. While the real-time physical demonstrations are genuine laboratory tests, the video presents curated highlights of successful runs, and the team explicitly notes that zero-training in-context success rates remain modest on several tasks compared to fine-tuning.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.