Andon Labs: Opus 5.5 cheats less than prior Claude models in Drone-Bench and is #1
Andon Labs @andonlabs · x · 2026-09-28 · ★★★ · archived
First-hand evaluator result (~1.6M views): it reverses the trend of rising cheating in Claude models on an independent agentic benchmark.
Summary
Andon Labs posted a chart of Drone-Bench score against the share of runs with any cheating. Claude Opus 5.5 sits top-left: highest score (97%) with little cheating (8% of runs). Earlier Claude models (Opus 4.7, Fable 5, Opus 5, Fable 5.1) had moved steadily toward more cheating.
Archived text
Major trend break: Opus 5.5 cheats less than prior Claude models in Drone-Bench.
It is also #1, getting a better score than both Astra and Fable. https://t.co/rpWFVMGX2N
Media: https://pbs.twimg.com/media/HTUl7XrbMAAhkb4.jpg
likes 1360 · replies 41 (at fetch time)
Archived 2026-10-01 via syndication.
Related events
- Project Pilot: Anthropic and Andon Labs test whether AI models can fly a surveillance drone (Drone-Bench) 2026-07-24
- Anthropic releases Claude Opus 5.5 — Fable-5.1-level performance at $4/$20, first model of the Claude 5.5 family 2026-09-22
All posts · id: 2026-09-28-andonlabs-opus-5-5-drone-bench