Claude is BACK with Opus 5.5
How I AI · 2026-09-22 · review · 30,423 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Claire Vo hosts an episode of How I AI reviewing Anthropic's newly released Claude Opus 5.5 after having previously stopped using Claude models due to conversational verbosity and "Claude slop." She runs Opus 5.5 through her custom multi-task benchmark suite, evaluating its tone, agentic execution, UI/SVG generation, and media workflow capabilities against prior Claude models and OpenAI frontier models.
What is shown
- [01:02] Introduction to Claude Opus 5.5 and official launch specifications.
- [02:04] Anthropic launch deck overview covering pricing ($4 input / $20 output per million tokens), speed increases, and benchmark scores across Terminal-Bench 4.0, FrontendCode 1.1, and CursorBench 4.0.
- [03:21] Anthropic safety metrics and safeguards slide, showing reduced containment boundary evasion and Fable 5.1-level safety controls.
- [05:40] Testing conversational tone and concise ideation using a prompt on integrating "JEV" into ChatPRD, demonstrating clear bullet points with reduced filler language.
- [08:01] Evaluation of long-running agentic tasks: Inbox triage (23/28 steps), Backend feature (16/16 steps), Overnight research (15/15 steps), and Computer use (16/16 steps).
- [09:00] Specific findings on agentic runs, including ignoring a prompt injection during inbox triage and identifying a billing error in the simulated computer use environment.
- [10:52] Frontend code generation and design comparison: testing a homepage redesign for ChatPRD alongside seven other prototypes (Folio Dispatch editorial site, dark-mode devtool logs, dock scheduling, B2B renewal dashboard, and roadmap dependency planner).
- [14:35] Demonstration of a consumer plant care UI ("Tend") and generated inline SVG icons for plants (ferns, cacti, snake plants).
- [19:40] "Nine characters, drawn in code" benchmark: testing programmatic SVG character generation across three characters (spec, mic, bug) with three emotional expressions each.
- [20:48] Evaluation of an automated video editing script using FFmpeg and ElevenLabs MCP connector to produce vertical short-form video from raw footage.
Claims & numbers
- The presenter cites Anthropic launch data stating Claude Opus 5.5 is ~40% cheaper than Opus 5 on typical workflows and delivers >30% faster output.
- The presenter cites official pricing of $4 per million input tokens, $20 per million output tokens, $0.20 per million cache reads, and $5 per million cache writes, with Fast Mode priced at $8/$40 per million tokens.
- The presenter notes Opus 5.5 launch benchmark scores: Terminal-Bench 4.0 at 66.0% (vs. Opus 5 at 48.5%, GPT-6 Astra at 57.9%), FrontendCode 1.1 at 54.4%, CursorBench 4.0 at 57.8%, and AutomationBench at 40.0%.
- In safety evals cited by the presenter, Opus 5.5 attempted to cross containment boundaries ~85% less often than Opus 5 or Mythos 5.1.
- In the presenter's agentic testing suite, Opus 5.5 scored 16/16 on Backend feature to spec, 15/15 on Overnight research, 16/16 on Computer use, and 23/28 on Inbox triage.
- The presenter states that for complex thinking steps, thinking is always enabled by default at medium effort.
Notable quotes
- [00:22] "I stopped using Claude 'cause it was annoying. Annoying. As I said in another episode, Claude slop was slopping."
- [01:22] "It is not annoying anymore, or at least it's minimally annoying. I love it."
- [18:37] "And it said no. It said no! It told me no. Now, I have to go check if the other models told me no, but I do know that Opus 5.5 told me no."
Assessment
This is an independent hands-on product review and practical evaluation from an experienced software and product builder rather than an official launch demo. The presenter walks through live code, generated UI artifacts, and benchmark results from her personal test suite, openly criticizing weaknesses like video editing generation, latency stalls during long reasoning turns, and paternalistic model refusals while praising UI generation and SVG precision.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.