NEW Sonnet 5.5 Is Opus 5 Level
Mehul Mohan · 2026-09-29 · community · 17,590 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Software engineer Mehul Mohan reviews Anthropic’s release of Claude Sonnet 5.5, analyzing its benchmark performance, pricing structure, and positioning within the Claude 5.5 family. He demonstrates using Sonnet 5.5 in Claude Code to implement a privacy toggle on his custom financial trading dashboard, highlighting the model's high coding speed alongside subtle instruction-following lapses compared to Claude Opus 5.5.
What is shown
- [00:00] Anthropic's announcement post on X detailing Claude Sonnet 5.5's release, speed improvements, and reduced token costs.
- [01:30] Anthropic's release blog post and launch documentation overview.
- [02:25] Evaluation benchmark table comparing Sonnet 5.5, Sonnet 5, Opus 5.5, and GPT-6 Sol across Terminal-Bench 4.0, FrontierCode 1.1, CursorBench 4.0, GDPval-AA, Humanity’s Last Exam, and OSWorld 2.1.
- [04:09] System card footnote detailing why Sonnet 5.5 scored lower at Max effort than Xhigh effort on FrontierCode due to Claude Code review subagent timeout issues.
- [05:37] Sponsor walkthrough of the Nebius Token Factory catalog and playground.
- [08:15] Cost and speed pricing table comparing Sonnet 5.5 ($2/$10 per 1M tokens) to Opus 5.5 ($4/$20 per 1M tokens) and cache read/write rates.
- [09:31] Side-by-side animated coding test generating an HTML/JS canvas simulation of a 400-starling murmuration.
- [10:35] Walkthrough of the presenter's personal Interactive Brokers trading dashboard.
- [12:54] Claude Code CLI terminal logs showing a prompt to add a privacy mode switch, Sonnet 5.5's unrendered implementation, and its subsequent fix after reviewing a user-submitted screenshot.
- [14:01] Whiteboard diagramming illustrating the "instruction following gap" between Opus 5.5 (100% completion) and Sonnet 5.5 (95% completion requiring manual correction).
- [16:03] Anthropic playbook article ("Building with Claude Sonnet 5.5" by Addy Osmani) outlining workload recommendations between Sonnet and Opus.
- [17:41] Mehul's post on X summarizing "sonnet is the new opus / opus is the new fable."
Claims & numbers
- The presenter highlights Anthropic's claim that Sonnet 5.5 runs more than 30% faster and costs up to 30% less for most tasks compared to Sonnet 5 [01:35].
- On Terminal-Bench 4.0 agentic coding, Sonnet 5.5 scores 70.6% compared to Sonnet 5 (10.3%) and Opus 5.5 (66.4%) [02:25].
- On FrontierCode 1.1, Sonnet 5.5 achieves 46.2% at Max effort and 52.1% at Xhigh effort, versus Sonnet 5 (42.4%), Opus 5.5 (54.4%), and GPT-6 Sol (49.3%) [02:26].
- On CursorBench 4.0, Sonnet 5.5 scores 55.5% versus 34.1% for Sonnet 5 and 57.8% for Opus 5.5 [02:26].
- Sonnet 5.5 scores 1844 on GDPval-AA v2.1 and 80.1% on OSWorld 2.1 computer use [03:49].
- Claude Sonnet 5.5 API pricing is set to $2 per million input tokens, $10 per million output tokens, $0.20 per million cache reads, and $2.50 per million cache writes, representing half the cost of Opus 5.5 on inputs, outputs, and cache writes [08:15, 08:48].
- The presenter estimates that 96% to 97% of heavy agentic token consumption consists of cache reads [08:40].
- The presenter claims Sonnet 5.5 typically achieves 95% of complex agentic tasks cleanly but regularly requires human intervention on the final 5%, whereas Opus 5.5 completes tasks with 100% reliability in his experience [14:18–14:45].
- The presenter notes OpenAI DevDay is scheduled for the following day with anticipated personal AI assistant announcements [18:08].
Notable quotes
- "Sonnet 5.5 scores more than Opus 5.5, which is a very, very interesting observation." [02:30]
- "Opus 5.5 is probably the best model ever... in the history of all AI models that I have personally used." [12:00]
- "Sonnet 5.5 is sort of like, it gets to 95%, right? You have to go ahead and push it at the rest of the 5%. With Opus 5.5, what I have seen is that this is happening at 100% every time." [14:18]
Assessment
An authentic developer review and hands-on appraisal. The presenter tests the model on real-world personal codebases via Claude Code and provides transparent terminal logs showing genuine errors and self-corrections alongside official benchmark comparisons.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.