Claude Sonnet 5 IS OUT & ITS HORRIBLE! Worst Model By Anthropic EVER? (Fully Tested)
WorldofAI · 2026-07-31 · review · 34,480 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
In this video, the presenter behind the YouTube channel WorldofAI reviews Anthropic's Claude Sonnet 5 model following its release. He analyzes its official benchmarks, pricing structure, and updated tokenizer, concluding that the model is inefficient and underwhelming compared to Claude Opus 4.8. He then tests Sonnet 5 on complex generation tasks, including an interactive macOS web clone, a voxel game, a SaaS landing page, and vector SVG art.
What is shown
- [00:00] Announcement and Agentic Gameplay: Displays Anthropic’s launch announcement and a gameplay capture of Claude Sonnet 5 playing a 3D space shooter using medium reasoning in a single shot.
- [00:27] Documentation and Benchmarks: Walks through Anthropic’s model comparison page, highlighting evaluations across SWE-bench Pro, Terminal-Bench 2.1, Humanity's Last Exam, and OSWorld.
- [01:28] Leaderboards and Pricing Fine Print: Explores the World of AI benchmark leaderboard and Anthropic's pricing documentation, pointing out footnote 2 regarding tokenizer density increases.
- [03:25] CursorBench & Token Efficiency: Displays leaderboard rankings showing Sonnet 5 Max at #13 and an Artificial Analysis chart plotting intelligence against output token consumption.
- [05:22] macOS Web Clone Test: An interactive browser-based macOS desktop simulation generated by Sonnet 5, complete with window management, a settings app, file browser, terminal, calculator, and an embedded raycaster FPS mini-game called Breach.
- [07:24] Minecraft Web Simulation Test: A 3D voxel sandbox in the browser with textured blocks, simple water physics, block placement, and basic mob renders (villager, creeper).
- [08:52] SaaS Landing Page Test: A landing page generated for an automated operations product ("Lumen"), demonstrating GSAP-style scroll triggers and layout bugs.
- [09:53] SVG Vehicle Generation: Side-by-side comparison of SVG renderings of a BMW M4 CS generated at various effort levels (low, medium, high).
Claims & numbers
- The presenter notes Anthropic's reported benchmark figures for Claude Sonnet 5:
- 63.2% on SWE-bench Pro (verified).
- 80.4% on Terminal-Bench 2.1.
- 43.2% on multidisciplinary reasoning (Humanity's Last Exam).
- 81.2% on computer use (OSWorld).
- 1,618 on knowledge work (GDPval AA v1.0).
- The presenter states introductory pricing is $2 per 1M input tokens and $10 per 1M output tokens through August 31, 2026, rising afterward to standard pricing of $3 input / $15 output per 1M tokens.
- The presenter states Sonnet 5 features a 1M token context window.
- The presenter highlights Anthropic's footnote showing that the new tokenizer (shared with Opus 4.7) maps text to roughly 1.0× to 1.35× more tokens depending on content type.
- On CursorBench, the presenter states Sonnet 5 Max ranks #13 scoring 61.2% at $6.87 (93,485 tokens), compared to Opus 4.8 Max at #8 scoring 63.8% at $7.59 (77,370 tokens)—making Sonnet 5 Max only $0.72 cheaper per task while burning more tokens.
- The presenter states the full macOS web desktop took approximately 40 minutes to generate in the workbench on Max mode.
Notable quotes
- [02:53]: "This means the same price of text can tokenize into roughly 1.0 times to 1.35 times more tokens than before, depending on the content."
- [03:52]: "It's only 72 cents cheaper than Opus 4.8 Max. At that point, it's defeating the purpose of just using the Sonnet model for everyday work..."
- [10:46]: "In conclusion, the Claude Sonnet 5 is totally underwhelming. I don't know what Anthropic was doing here, and it is something that you should not use at all."
Assessment
This is a critical third-party product review and hands-on benchmark evaluation, not an official launch video. The presenter tests real code outputs and compares published API pricing and tokenization metrics, highlighting practical inefficiencies that contrast with initial launch marketing.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.