RoboArena: crowd-sourced, double-blind real-world evaluation of generalist robot policies
RoboArena (arXiv 2506.18123, 2025-06-22) ranks generalist robot policies through double-blind pairwise comparisons run by a distributed network of evaluators on the DROID platform, who pick their own tasks and scenes. The first round covered 600+ real-robot episodes over 7 policies at 7 academic institutions; its open leaderboard became a standard reference, e.g. NVIDIA's GR00T N2 and Cosmos 3 claims in 2026.
Key facts
- Paper: 'RoboArena: Distributed Real-World Evaluation of Generalist Robot Policies' (Atreya, Pertsch, Lee, Kim et al.), arXiv 2506.18123; published at CoRL 2025 (PMLR v305)
- 612 pairwise real-robot comparisons, 7 generalist policies, 7 universities, DROID Franka setup
- Authors show this ranks policies more accurately than centralized fixed-task evaluation
- Evaluation network opened to the community
What happened
RoboArena borrowed the idea behind Chatbot Arena, pairwise preference votes aggregated into a ranking, and applied it to physical robots. Evaluators at partner universities run two anonymous policies on a task of their choice and record which did better.
Why it matters
Real-world robot evaluation is expensive and hard to standardize. A distributed arena gives a scalable, harder-to-game ranking of VLAs, and labs now cite it in model launches.
Changelog
- 2026-09-29: created (author affiliations not verified; the paper lists Atreya, Pertsch, Lee, Kim among the authors)
Related events
- Ai2 launches MolmoSpaces, an open simulation ecosystem and leaderboard for generalist robot policies ★★
- NVIDIA releases Cosmos 3, an open omni-model for physical AI (world generation, reasoning and actions) ★★★
Sources (2)
id: 2025-06-22-roboarena · updated 2026-09-29 · open in the interactive timeline