I Tested Sonnet 5.5 vs Opus 5.5 (WILD RESULTS)
Brock Mesarich | AI for Non Techies · 2026-09-28 · review · 48,104 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
An independent presenter evaluates and benchmarks Anthropic’s Claude Sonnet 5, Sonnet 5.5, Opus 5.5, and Fable 5.1 by having each model generate a full 3D interactive browser game from an identical detailed prompt. He tests the playable outputs in real-time, assessing gameplay, visual quality, and stability while tracking the total generation time and API cost for each model.
What is shown
- [00:15] Scorecard overview on Excalidraw comparing Sonnet 5, Sonnet 5.5, Opus 5.5, and Fable 5.1.
- [00:45] Pricing breakdown table comparing Claude Sonnet 5.5 and Claude Opus 5.5 per 1 million tokens.
- [01:34] The benchmark prompt detailing constraints (single self-contained
index.html, procedural geometry/shaders, 60fps, day/night cycle, weather, audio, interaction). - [02:00] Playtesting Sonnet 5's output ("The Edge of the Sky"), showing glitchy water, simple geometry, and limited interaction.
- [03:04] Sonnet 5 results recorded: 7:06 generation time, $1.59 cost.
- [03:16] Playtesting Opus 5.5's output ("Aerie"), showing detailed terrain, water surface effects, physics-based rock throwing, and dynamic fog.
- [04:26] Opus 5.5 results recorded: 41:47 generation time, $11.95 cost.
- [04:51] Playtesting Fable 5.1's output ("Aerie"), featuring ancient ruins and interactive elements, but accompanied by screen-shaking movement glitches.
- [05:56] Fable 5.1 results recorded: 40:30 generation time, $17.45 cost.
- [06:28] Playtesting Sonnet 5.5's output ("Skyreach"), demonstrating animated hopping rabbits, procedural grass, smooth movement, swimming fish, and dynamic weather/rain.
- [07:27] Sonnet 5.5 results recorded: 36:43 generation time, $9.02 cost.
- [08:01] Side-by-side visual comparison and final scorecard review across all four models.
Claims & numbers
- Anthropic official release claims cited by presenter: Claude Sonnet 5.5 runs 30%+ faster, costs up to 30% less for most work, and requires fewer tokens per task than Sonnet 5 [00:36, 01:18].
- Pricing cited per 1M tokens [00:54]:
- Claude Sonnet 5.5: Cache reads $0.20, Cache writes $2.50, Input tokens $2.00, Output tokens $10.00.
- Claude Opus 5.5: Cache reads $0.20, Cache writes $5.00, Input tokens $4.00, Output tokens $20.00.
- Benchmark test results (run on "effort level: high"):
- Sonnet 5: 7 minutes 6 seconds; $1.59.
- Sonnet 5.5: 36 minutes 43 seconds; $9.02.
- Opus 5.5: 41 minutes 47 seconds; $11.95.
- Fable 5.1: 40 minutes 30 seconds; $17.45.
Notable quotes
- [00:40] "It runs 30% faster and costs up to 30% less for most of the work."
- [04:28] "So, Opus 5.5 costed, drum roll please, $11.95."
- [06:38] "I personally think this might be the most, like the best looking world."
Assessment
This is an authentic third-party benchmark and hands-on comparison demonstrating the execution of code generated by different LLMs. The generation processes took place prior to recording, but the presenter plays the unedited resulting web games directly in Chrome and displays exact recorded generation durations and API costs.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.