Getting the most out of Opus 5.5
Theo - t3․gg · 2026-09-24 · tutorial · 218,471 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Theo Browne (t3.gg) reviews best practices for using Anthropic’s Claude Opus 5.5 in Claude apps and Claude Code, walking through an official playbook written by Addy Osmani. Throughout the video, Theo tests agent workflows in his T3 Code environment, analyzes benchmark data comparing reasoning levels and model code-review quality, and explains how to properly steer long-running autonomous coding runs.
What is shown
- [02:24] Addy Osmani’s playbook article titled "Getting the most out of Opus 5.5 in Claude and Claude Code".
- [04:15] Demonstrating a long-running T3 Code session executing autonomous refactoring on a remote laptop ("LeftBook"), highlighting prompts specifying "done" states, explicit environment permissions, and instructions to ask questions when blocked.
- [10:38] Reviewing the playbook's guidance on defining clear exit criteria ("done" states) for tasks rather than open-ended objectives.
- [12:41] A benchmark spreadsheet ("skatebench") analyzing Claude Opus 5.5 reasoning levels (
xhighvs.max), comparing token counts, durations, and accuracy. - [16:26] The Which AI Made This? interface comparing frontend UI design outputs between Claude Fable 5.1 and Claude Opus 5.5.
- [18:48] Editing a local
CLAUDE.mdrules file in VS Code to include steering instructions on when to continue working autonomously versus when to stop and ask for human confirmation. - [19:24] Dispatching updated repository rules across an agent fleet using Claude Opus 5.5 in T3 Code.
- [21:09] Reviewing guidelines for inspecting agent final summaries and asking models to evaluate rollout risks and review code diffs.
- [23:16] A benchmark scorecard measuring confirmed code issue findings across models (GPT-6 Astra, Grok 4.7, GPT-6 Sol, Fable 5.1, Claude Opus 5.5, Opus 5, and Gemini 3.8 Flash High).
- [25:54] Examining Claude app safety mechanisms, auto-model downgrades upon safety flags, and settings to disable automatic switching.
Claims & numbers
- The presenter states that Addy Osmani, previously on Google's Chrome team, recently joined Anthropic (article published September 22, 2026) [00:26].
- On the Skatebench benchmark, the presenter claims Opus 5.5 on
xhighaveraged 338 tokens per response and a 6-second average duration (slowest response: 31 seconds) [13:05]. - On
maxreasoning in Skatebench, the presenter states Opus 5.5 average tokens increased over 10x to 5,000, average duration rose to 50 seconds, and the slowest run hit 600 seconds, while benchmark accuracy only increased from 78% to 79% (costing 13x more and using 15x tokens for one additional correct answer) [13:14]. - The presenter claims
maxreasoning does not make models smarter, but forces them not to think less by removing their ability to stop reasoning early [12:31]. - In a code audit benchmark on the T3 Code repository shown on screen:
- GPT-6 Astra scored 83.8 confirmed quality (8 supported findings) [23:40].
- Grok 4.7 scored 80.7 (8 supported findings) [23:31].
- GPT-6 Sol scored 79.9 (9 supported findings) [23:47].
- Claude Fable 5.1 scored 69.7 (5 supported findings) [23:55].
- Claude Opus 5.5 scored 67.5 (5 supported findings, zero contradicted/unresolved) [24:12].
- Older Claude Opus 5 scored 37.9 (4 supported findings, 2 unconfirmed/contradicted) [24:20].
- The presenter notes that Opus 5.5 is the first Opus model to ship with Fable-level bio and cyber safety filters [25:57].
Notable quotes
- [12:31] "Max isn't just making it so the model can think more, it is removing its ability to think less."
- [14:47] "Don't tell the model to fucking think, it knows that it should think. It is smarter than you probably think."
- [28:38] "Also, do not touch max mode. Seriously, it's so bad."
Assessment
This is an authentic hands-on technical review and tutorial evaluating Claude Opus 5.5 and official Anthropic prompt-engineering recommendations. The presenter demonstrates live and recent local agent runs, shares real benchmark data from internal tests, and provides critical analysis of model behaviors without deceptive staging.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.