MolmoAct 2 / MolmoAct 2-Think
Checkpoints: MolmoAct2 (post-trained multi-embodiment foundation, ~5.4B params per HF safetensors), -Think, -Pretrain, fine-tuned -DROID, -BimanualYAM, -SO100_101, -LIBERO, -Think-LIBERO, FAST-Tokenizer. Main supported robots: SO-100/101, bimanual YAM, Franka (DROID); others need fine-tuning. Paper arXiv 2605.02881.
- Input
- text, image
- Output
- action
- License
- Apache-2.0 (code); model weights on HF (license tag not stated on card)
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Hugging Face | allenai/MolmoAct2 | huggingface.co/collections/allenai/molmoact2-models | — |
| GitHub | — | github.com/allenai/molmoact2 | — |
| Hugging Face LeRobot | allenai/MolmoAct2-LIBERO-LeRobot | huggingface.co/allenai/MolmoAct2-LIBERO-LeRobot | — |
Notable capabilities (4)
- Open action reasoning model: Molmo2-ER embodied-reasoning VLM connected to a flow-matching action expert via per-layer KV conditioning; the Think variant adds adaptive depth reasoning (interpretable depth map before acting). source
- Strong out-of-the-box real-world success: 87.1% average success over 15 real Franka tasks vs 45.2% for π0.5 and 48.4% for MolmoBot (Ai2's evaluation); LIBERO 97.2% (98.1% Think). source
- Fast inference: ~180 ms per action call (790 ms with adaptive depth reasoning) vs ~6,700 ms for the original MolmoAct (up to 37x faster). source
- Largest open bimanual dataset: Released with MolmoAct2-BimanualYAM, 720+ hours of bimanual tabletop demonstrations, which Ai2 calls the largest open bimanual robotics dataset, plus an open FAST action tokenizer. source
Fully open (weights, data, code) VLA from Ai2, the main open alternative to π0.5/GR00T for tabletop manipulation. Start from a fine-tuned checkpoint (e.g. allenai/MolmoAct2-DROID) for ready-to-run inference; the base card has no inference code.
Sources: Ai2 blog, arXiv 2605.02881, HF model card, GitHub.
Timeline entry
- Ai2 releases MolmoAct 2, a fully open robot action-reasoning model that beats π0.5 on real-world tasks ★★
On 2026-05-05 the Allen Institute for AI released MolmoAct 2 and MolmoAct 2-Think, open vision-language-action models built on the Molmo2-ER embodied-reasoning VLM with a flow-matching action expert, along with weights, code and 720+ hours of bimanual data. In Ai2's tests it reached 87.1% average…