Reka unveils Rho-1, a 19B 'omni' model that reads and generates text, images, streaming video and robot actions in one context
On Oct 5, 2026 Reka published Rho-1, a 19B-parameter model trained from scratch that puts text, vision, video and robot actions as tokens in a single context, streaming steerable video at about 0.79x real time; a distilled Flash variant makes a 5.3-second clip in about one second. It is a research preview only: no public weights or API.
Key facts
- 19B parameters, 'symmetric two-stream transformer with shared attention', trained from scratch on 320 H100 GPUs for three months (Reka)
- Streams video at a median 0.79x real time (99 denoising steps); a distilled 8-step variant (Rho-1 Flash) renders a 5.3-second clip in ~1 second
- Native video resolution 672x384; continuous video can be steered mid-rollout without cuts
- Robot actions: demo episode on LIBERO simulation tasks (seven action channels); no quantitative LIBERO scores published. Companion inverse-dynamics model derives robot controls from ordinary internet video
- Known limits stated: long-horizon drift in 30-second streams, unreliable object tracking, brittle editing
- Availability: research preview through collaboration on Reka Cloud; no public weights or public API (contact@reka.ai)
What happened
Reka, the small multimodal lab behind the Reka Core/Flash/Edge models, published a research post on Rho-1. It is a single 19B transformer that handles text, images, video and robot actions as tokens in one context window, instead of chaining a language model with a separate video diffusion model. Reka says it is "among the fastest models in the world in every modality" and "faster than any other video model we timed" for 5.3-second clips, but it names no competing models. Every result comes from a checkpoint trained on "just 320 H100 GPUs for three months".
Why it matters
This is a compute-light attempt at the unified "omni" world-model and robotics stack that larger labs pursue with far more compute. The speed claims are Reka's own. No independent benchmarks, weights or API exist yet.
Changelog
- 2026-10-05: created (21:30 full run)
Sources (3)
- officialReka: Rho-1: Collapsing the multimodal stack
- officialReka: inverse dynamics model for interactive world models
- pressAlphaSignal: Reka's Rho-1 merges reasoning, video, and robot control into one model
id: 2026-10-05-reka-rho-1-omni-model · updated 2026-10-05 · open in the interactive timeline