I Tested Sonnet 5.5 (Here Is What You Need to Know)
Никита Ефимов | ИИ и автоматизация · 2026-09-29 · review · 9,171 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Nikita Efimov reviews Anthropic's newly released Claude Sonnet 5.5, analyzing its benchmark performance, pricing structure, and cost efficiency relative to Claude Opus 5.5 and Claude Fable 5.1. He demonstrates why Sonnet 5.5's 50% cheaper token price does not necessarily translate to lower task costs on complex agentic workflows due to the model's higher token consumption at elevated "effort" settings. Efimov provides practical workflow recommendations, suggesting Sonnet 5.5 for lightweight daily routines and Opus 5.5 for demanding engineering and reasoning tasks.
What is shown
- [00:53] A pixel-art animated intro video created with Claude Sonnet 5.5, depicting the recent sequence of releases (Opus 5.5, GPT-6 Sol and Luna, Sonnet 5.5).
- [02:05] Anthropic's model tier overview table (Claude Fable 5.1, Opus 5.5, Sonnet 5.5, Haiku 4.5) displaying pricing per million input/output tokens and model roles.
- [03:18] Official Anthropic benchmark comparison table covering Terminal-Bench 4.0, FrontierCode 1.1, CursorBench 4.0, GDPval-AA v2.1, AA Briefcase 1.4, Humanity’s Last Exam, and OSWorld 2.1.
- [04:29] Anthropic’s thinking "Effort" settings (Low, Medium, High, Extra High, Max) and an explanation of their relationship to token expenditure.
- [05:00] Benchmark accuracy vs. cost curves for Terminal-Bench 4.0 and CursorBench 4.0 comparing Sonnet 5.5, Opus 5.5, and GPT-6 Sol across different effort settings.
- [06:53] FrontierCode 1.1 accuracy vs. cost chart showing Sonnet 5.5's performance degradation and cost surge at the "Max" effort setting.
- [09:36] Claude.ai free tier interface and capabilities overview (web search, file uploads, artifacts, projects).
- [10:48] Chat window best practices diagram ("1 task = 1 chat" to preserve rate limits).
- [11:39] Diagram of Efimov's updated multi-model workflow architecture (Opus 5.5 for planning and deep reasoning; Sonnet 5.5 for routine tasks).
- [12:57] Anthropic optimization documentation discussing single-model vs. multi-model agent pipeline costs.
- [14:06] Claude Code settings showing permission modes (Plan, Accept Edits, Auto, Bypass Permissions).
- [14:28] Excerpt from Anthropic's Sonnet 5.5 System Card describing cybersecurity safeguards and automatic fallback to Sonnet 5.
Claims & numbers
- Release pacing: The presenter states that three major models launched within one week: Opus 5.5 on September 22, GPT-6 Sol 1.5 hours later, and Sonnet 5.5 on September 28 [00:09].
- API pricing: Sonnet 5.5 is priced at $2/MTok input and $10/MTok output—exactly half the price of Opus 5.5 ($4/$20 MTok), while Haiku 4.5 is $1/$5 MTok and Fable 5.1 is $10/$50 MTok [02:08, 02:45].
- Speed: Anthropic claims Sonnet 5.5 is 30% faster than Sonnet 5 [02:50].
- Terminal-Bench 4.0 scores: Sonnet 5.5 scored 70.6%, beating Opus 5.5 (66.4%) and significantly surpassing Sonnet 5 (10.3%) [03:40].
- Benchmark gaps: Sonnet 5.5 trails Opus 5.5 by only 1–3% on several evaluations: CursorBench 4.0 (55.5% vs. 57.8%), OSWorld 2.1 (50.1% vs. 81.8%), and Humanity's Last Exam (64.5% vs. 67.7%) [04:05].
- Terminal-Bench effort/cost comparison: At "Extra High" effort, Sonnet 5.5 scores 61.5% at a cost of $5.30 per attempt; Opus 5.5 at standard "High" effort scores 64.2% at $3.88 per attempt [05:08].
- CursorBench effort/cost comparison: Sonnet 5.5 at "Max" effort reaches 55.5% accuracy at $9.67 per task, while Opus 5.5 at "High" effort scores 56.0% at $3.97 per task [05:42].
- Performance drop at Max effort: On FrontierCode 1.1, increasing Sonnet 5.5 effort to "Max" drops accuracy from 52.1% (Extra High) to 46.2%, while attempt cost surges 13x from $1.59 to $20.78 due to overthinking and excessive self-verification loops [06:56].
- Low-effort cost: Simple routine tasks run on Sonnet 5.5 at "Low" effort cost between $0.20 and $0.80 per task via API [08:50].
- Document tasks: At "Low" effort, Sonnet 5.5 matches Opus 5.5 output quality while being approximately 25% cheaper [09:03].
- Subscription tiers: Sonnet 5.5 is available on Claude.ai's free tier (with 5-hour rate-limit resets), while the Pro tier ($20/month) offers 5x higher message limits and access to Opus 5.5 and Claude Code [09:36, 11:03].
- Anthropic single vs. multi-model study: Anthropic docs show that a single model at lower effort is cheaper than chaining two models (e.g., Opus 5.5 alone at High costs $1.38 vs. Opus 5.5 with a Fable 5.1 advisor at $2.92) [13:04].
- Cybersecurity guardrails: High-risk cybersecurity prompts trigger automatic fallback from Sonnet 5.5 to Sonnet 5 [14:35].
Notable quotes
- [05:27] "То есть Opus и умнее, и дешевле." ("That is, Opus is both smarter and cheaper.")
- [06:01] "Потому что в два раза дешевле у него слово, а задача выходит столько же." ("Because its price per word is twice as cheap, but the whole task costs the same.")
- [07:05] "На максимуме модель начинает перестраховываться. Он запускает кучу ненужных проверок..." ("At maximum, the model starts over-insuring itself. It runs a bunch of unnecessary checks...")
Assessment
This is an independent software review and strategy breakdown evaluating the real-world utility of Anthropic's Claude Sonnet 5.5 release. The host does not perform live coding on camera, relying instead on official benchmark charts, system cards, and documented pricing data to argue convincingly that token-level discounts do not always translate to cheaper task execution.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.