Anthropic's 'Claude plays robotics': LLMs fail at direct joint control but succeed when supervising controllers
On July 9, 2026 Anthropic published "Claude plays robotics", which tested twelve models from five providers on simulated and physical robots. Models mostly fail when they must drive joints directly, with 0–5.5% direct manipulation success (best: Claude Mythos Preview). They complete real navigation and manipulation tasks when they supervise a pretrained controller or a VLA policy, or use simple tools such as a compass.
Key facts
- Authors: Shmuel Berman, Michael Ilie, Jia Deng, Daniel Freeman
- 12 models from 5 providers, several Claude generations
- Direct manipulation success 0–5.5%; Claude Mythos Preview highest at 5.5%
- Quadruped: newer models reach 'limited but meaningful' whole-body control (balancing almost two seconds); no model stood a collapsed humanoid up
- VLA supervision on LIBERO-40 raised success substantially for all models
- A compass tool beat visual self-views; extended reasoning helped little and sometimes hurt
What happened
A systematic eval of general LLMs as robot controllers at several levels of abstraction, from raw joint commands up to supervising pretrained policies.
Why it matters
The study measures how far general frontier models are from embodied control. The answer depends heavily on the interface: high-level supervision already works, while low-level control improves only now and then from one generation to the next.
Changelog
- 2026-10-01: created (leads run, from the Anthropic uncited-posts audit)
Related events
- Project Pilot: Anthropic and Andon Labs test whether AI models can fly a surveillance drone (Drone-Bench) ★★★
- Anthropic's robot exposure index: robots can technically do 74% of physical job tasks but are cost-competitive on only 0.3% ★★★
Sources (1)
id: 2026-07-09-anthropic-claude-plays-robotics · updated 2026-10-01 · open in the interactive timeline