Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. Anthropic's 'Claude plays robotics': LLMs fail at direct…

Anthropic's 'Claude plays robotics': LLMs fail at direct joint control but succeed when supervising controllers

★★after cutoffroboticsAnthropicconfidence: high

On July 9, 2026 Anthropic published "Claude plays robotics", which tested twelve models from five providers on simulated and physical robots. Models mostly fail when they must drive joints directly, with 0–5.5% direct manipulation success (best: Claude Mythos Preview). They complete real navigation and manipulation tasks when they supervise a pretrained controller or a VLA policy, or use simple tools such as a compass.

Key facts

What happened

A systematic eval of general LLMs as robot controllers at several levels of abstraction, from raw joint commands up to supervising pretrained policies.

Why it matters

The study measures how far general frontier models are from embodied control. The answer depends heavily on the interface: high-level supervision already works, while low-level control improves only now and then from one generation to the next.

Changelog

  • 2026-10-01: created (leads run, from the Anthropic uncited-posts audit)

Related events

  1. Project Pilot: Anthropic and Andon Labs test whether AI models can fly a surveillance drone (Drone-Bench) ★★★
  2. Anthropic's robot exposure index: robots can technically do 74% of physical job tasks but are cost-competitive on only 0.3% ★★★

Sources (1)

id: 2026-07-09-anthropic-claude-plays-robotics · updated 2026-10-01 · open in the interactive timeline