Cosmos 3 (Nano / Super)
Announced at GTC 2026-03-16 ('the first world foundation model unifying synthetic world generation, vision reasoning and action simulation' - NVIDIA claim); weights published 2026-05-31/06-01 (HF blog 'The First Open Omni-model for Physical AI Reasoning and Action'). Sizes: Nano 16B, Super 64B. Linux + Ampere/Hopper/Blackwell GPUs, BF16. Technical report dated 2026-06-22.
- Input
- text, image, video, audio, action
- Output
- text, image, video, audio, action
- License
- OpenMDW-1.1 (commercial and non-commercial use)
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Hugging Face | nvidia/Cosmos3-Nano | huggingface.co/nvidia/Cosmos3-Nano | — |
| Hugging Face | nvidia/Cosmos3-Super | huggingface.co/nvidia/Cosmos3-Super | — |
| GitHub | — | github.com/nvidia-cosmos | docs |
Notable capabilities (3)
- FIRST Unified omni world model (generation + reasoning + action): One Mixture-of-Transformers model (autoregressive + diffusion) replaces separate Cosmos Predict, Transfer, Reason and Policy models: world generation, physical reasoning, forward/inverse dynamics and action/policy generation. source
- Open omnimodal I/O: Inputs text, images, short video, audio and action trajectories (16-400 frames); outputs text, images, video (5-400 frames), 48 kHz stereo audio and actions (JSON). source
- Leaderboard results (found after launch): NVIDIA cites best open text-to-image and image-to-video models on Artificial Analysis and best policy model on RoboArena. source
Sources: HF blog, Cosmos3-Nano, Cosmos3-Super, technical report.
Timeline entry
- NVIDIA releases Cosmos 3, an open omni-model for physical AI (world generation, reasoning and actions) ★★★
NVIDIA published open weights for Cosmos 3 (Nano 16B, Super 64B) around 2026-06-01: one Mixture-of-Transformers model that takes text, images, video, audio and robot actions and generates video, images, audio, text or actions, replacing the separate Cosmos Predict, Transfer, Reason and Policy…
Other NVIDIA models
NVIDIA Nemotron 3.5 Lightning (30B-A3B) · NVIDIA NemotronLabs VoiceChat 11B (and PersonaPlex-7B) · NVIDIA Nemotron 3 Ultra (550B-A55B) · NVIDIA Nemotron 3 Nano Omni (30B-A3B Reasoning) · Isaac GR00T N1.7 · NVIDIA Nemotron 3 Super (120B-A12B) · Cosmos Reason 2 · NVIDIA MagpieTTS Multilingual 357M · NVIDIA Parakeet / Canary / Nemotron Speech ASR (open) · Isaac GR00T N2 · Cosmos Predict 2.5 / Transfer 2.5 · Isaac GR00T N1 / N1.5 / N1.6