Claude Sonnet 5.5 - Benchmarks and Pricing | Beats Opus 5.5 and GPT-6 Sol?
United Top Tech · 2026-09-29 · review · 8,297 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
This video is a review presented by the creator of the channel United Top Tech covering Anthropic's launch of Claude Sonnet 5.5. The presenter walks through Anthropic's official announcement post, benchmark comparisons against Sonnet 5, Opus 5.5, and GPT-6 Sol, web UI availability on the free tier, and API pricing documentation.
What is shown
- [00:00] Anthropic's announcement post on X (@claudeai) introducing Claude Sonnet 5.5 as the second model in the Claude 5.5 family.
- [00:26] Official benchmark scorecard comparing Sonnet 5.5 against Sonnet 5, Opus 5.5, and GPT-6 Sol across agentic coding (TerminalBench, FrontierCode 1.0, CursorBench 4.0), knowledge work (AIA-Briefcase 1.1 and 1.0), multidisciplinary reasoning (Humanity's Last Exam), computer use (OSWorld 1.1), and visual chart recognition.
- [01:59] An X chart showing "Knowledge work by effort level (AIA-Briefcase 1.1)", tracking score versus cost per task for Sonnet 5.5, Opus 5.5, Sonnet 5, and GPT-6 Sol.
- [02:10] Anthropic's post confirming Claude Sonnet 5.5 is available immediately and teasing Claude Haiku 5.5 in upcoming weeks.
- [02:17] Claude web interface under a free plan, showcasing the model picker dropdown with Sonnet 5.5, effort level configurations (Low, Medium, High, Extra, Max), and adjacent options (Claude Fable 5.1, Opus 5.5, Haiku 4.5).
- [02:27] Anthropic Platform Documentation model comparison table displaying comparative latency, context window, and token pricing for Claude Fable 5.1, Opus 5.5, Sonnet 5.5, and Haiku 4.5.
Claims & numbers
- Speed and Cost: The presenter and official post state Sonnet 5.5 runs over 30% faster and costs up to 30% less for most work compared to Sonnet 5.
- Coding Benchmarks: On TerminalBench agentic coding, Sonnet 5.5 scores 70.6% versus Sonnet 5's 10.3% and Opus 5.5's 66.4%. On CursorBench 4.0, Sonnet 5.5 scores 55.0% versus Sonnet 5's 34.1% and Opus 5.5's 57.8%. On FrontierCode 1.0 (dev), Sonnet 5.5 reaches 46.2% (and 52.9% at high effort) compared to Opus 5.5's 54.4% and GPT-6 Sol's 49.3%.
- Knowledge Work: On AIA-Briefcase 1.1, Sonnet 5.5 scores 1844, matching Opus 5.5 (1844) and beating Sonnet 5 (1449) and GPT-6 Sol (1483). On AIA-Briefcase 1.0, Sonnet 5.5 scores 1811 versus Opus 5.5's 1822.
- Reasoning and Vision: On Humanity's Last Exam (with tools), Sonnet 5.5 reaches 64.5% compared to Opus 5.5's 67.7% and Sonnet 5's 54.9%. On visual chart recognition (ChartQA), Sonnet 5.5 scores 61.6% versus Opus 5.5's 64.4% and GPT-6 Sol's 52.6%.
- Pricing: The presenter highlights that Sonnet 5.5 costs $2 / million input tokens and $10 / million output tokens, half the price of Opus 5.5 ($4 / input, $20 / output).
Notable quotes
- [00:10] "It almost cooks the Opus 5.5 model, which is one of the top models in the world."
- [01:19] "That's a crazy jump."
- [02:39] "So it's almost half the price, and it gives this staggering benchmarks."
Assessment
This is a tech commentary and reaction video summarizing Anthropic's public announcement, documentation, and benchmark tables for Claude Sonnet 5.5. The presenter does not run independent evaluations or live benchmarks during the video, relying instead on official Anthropic documentation and X posts while demonstrating that the model is accessible in the free web interface.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.