NEW Claude Sonnet 5 vs Opus 4.8! (Full Review)
AI Foundations · 2026-07-31 · review · 48,700 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Drake from AI Foundations reviews Anthropic's newly released Claude Sonnet 5, comparing its benchmark results, pricing, and agentic coding capabilities directly against Claude Opus 4.8 and Claude Sonnet 4.6. He pits Sonnet 5 against Opus 4.8 side by side inside Claude Code using the /goal command to build an interactive canvas browser game called "Orbit Runner," evaluating speed, token usage, gameplay mechanics, and overall project cost.
What is shown
- [00:00 - 03:40] Official Anthropic announcement page for Claude Sonnet 5 (dated June 30, 2026), detailing model descriptions, benchmark comparisons against Sonnet 4.6 and Opus 4.8, and API pricing tables.
- [04:08 - 05:57] Side-by-side terminal setup in Claude Code comparing Opus 4.8 (left) and Sonnet 5 (right), both set to "Extra" effort level, receiving identical prompt specifications to build a single-page canvas game called "Orbit Runner."
- [05:58 - 09:15] Execution comparison: Sonnet 5 immediately initializes npm, installs Playwright, and writes automated tests while Opus 4.8 spends extensive time in internal reasoning before generating code. Opus 4.8 finishes in 9.8k tokens, while Sonnet 5 uses 13k tokens while running Playwright headless browser checks.
- [09:40 - 11:45] Side-by-side playtesting of the two generated games running on localhost; Drake plays both versions, showing differences in physics, UI styling, and directional thrust indicators before revealing which model generated each.
- [12:44 - 14:40] Drake prompts Claude Code to calculate the exact cost differences between the runs based on API token pricing.
- [15:33 - 16:16] Drake demonstrates his local autonomous workflow directory (
ai-foundations), showcasing eight custom skill modules across marketing, sales, and product management that can be transitioned from Opus 4.8 to Sonnet 5.
Claims & numbers
- SWE-bench Pro: The presenter shows Sonnet 5 scoring 63.2%, compared to 58.1% for Sonnet 4.6 and 69.2% for Opus 4.8 [00:54].
- Terminal-Bench 2.1: Sonnet 5 scored 80.4%, Sonnet 4.6 scored 67.0%, and Opus 4.8 scored 82.7% [01:38].
- Humanity's Last Exam (Multidisciplinary Reasoning): Sonnet 5 scored 43.2% without tools and 57.4% with tools, compared to Opus 4.8 at 49.8% without tools and 57.9% with tools [01:57].
- OSWorld Verified (Computer Use): Sonnet 5 scored 81.2%, Sonnet 4.6 scored 78.5%, and Opus 4.8 scored 83.4% [02:34].
- GDPval-AA v2 (Knowledge Work): Sonnet 5 scored 1618, higher than both Sonnet 4.6 (1395) and Opus 4.8 (1615) [02:44].
- Pricing: The presenter notes Sonnet 5 launched with introductory pricing of $2 per million input tokens and $10 per million output tokens through August 31, 2026, shifting to standard pricing of $3 per million input and $15 per million output tokens. Opus 4.8 regular pricing is $5 per million input and $25 per million output tokens [02:53, 03:19].
- Experiment Cost Comparison: Building the game cost approximately $0.13 in output tokens with Sonnet 5 (13,000 tokens used), compared to $0.245 with Opus 4.8 (9,800 tokens used) [13:50, 14:03].
Notable quotes
- "This is like a no-brainer. You're going to be saving so much money when using Claude Sonnet and sacrificing very little quality." [00:24]
- "Sonnet 5 is flying. It's already running tasks, installing projects, while Opus is taking a different strategy. Opus is thinking through this task a lot more." [06:01]
- "You get the same level of quality pretty much for half the cost, and I think a 50% price decrease is worth the quality in this one test that I did." [15:13]
Assessment
This is an authentic hands-on review and head-to-head coding benchmark by an independent creator testing newly released models via Claude Code. The creator shows unedited terminal outputs, realistic token/cost calculations, and directly playable localhost game implementations without staged effects.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.