Claude Sonnet 5 just dropped. I'm changing how I use AI...
Alex Finn · 2026-07-31 · community · 56,620 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Alex Finn reviews Anthropic's newly released Claude Sonnet 5, evaluating its benchmark performance, pricing, and agentic coding capabilities. He compares its 3D graphics generation against ChatGPT 5.5, outlines a cost-saving hybrid workflow pairing Claude Opus 4.8 for planning with Sonnet 5 for execution, and examines leaked strings indicating an impending return of Claude Fable 5.
What is shown
- Benchmark & Cost Breakdown [00:41, 01:29, 03:14]: Slides comparing Claude Sonnet 5 against Sonnet 4.6 and Opus 4.8 across SWE-bench Verified, Terminal-Bench 2.1, Humanity's Last Exam, OSWorld, and BrowseComp. Alex also shows his personal Hermes API billing dashboard displaying $1,375.94 in Claude usage over the previous month [01:58].
- Head-to-Head 3D Simulation Test [04:16]: A prompt requesting a single-file Three.js stormy sea simulation with a wooden sailing ship, Gerstner waves, dynamic lighting, rain, and UI controls is submitted to both Claude Code Desktop (running Sonnet 5) and OpenAI Codex (running ChatGPT 5.5).
- Output Inspection [04:52]: The ChatGPT 5.5 generation renders rain particles and controls, but features static ocean meshes and a stationary boat without camera rotation. In contrast, Sonnet 5's generation [05:27] produces an interactive 3D scene with dynamic wave physics, a rolling and pitching ship, and responsive controls.
- Hybrid Planning Workflow Demo [06:42]: In Claude Code Desktop, Alex sets Plan Mode to Opus 4.8 in "Ultra Code" mode to design an AI-powered Notion clone, launching five sub-agents in a background workflow [08:31]. Once the architectural markdown plan is generated [08:52], he switches the model to Sonnet 5 (Medium) to execute the implementation cheaply [09:08].
- Hermes Agent Setup & Fable 5 Leak [09:30, 10:15]: Switching the model selector in Hermes Agent/OpenClaw to Sonnet 5 via API, followed by a review of leaked Claude Code strings indicating upcoming API billing and identity verification requirements for Claude Fable 5.
Claims & numbers
- Benchmarks (Sonnet 5 vs Sonnet 4.6 vs Opus 4.8):
- SWE-bench Verified (Agentic coding): Sonnet 5 scores 63.2% vs Sonnet 4.6 at 58.1% and Opus 4.8 at 69.2% [02:45].
- Terminal-Bench 2.1 (Agentic coding): Sonnet 5 scores 80.4% vs Sonnet 4.6 at 67.0% and Opus 4.8 at 82.7% [02:45].
- Humanity's Last Exam (Multidisciplinary reasoning): Sonnet 5 scores 43.2% with vision / 57.4% text-only vs Sonnet 4.6 at 34.6% / 46.8% and Opus 4.8 at 49.8% / 57.9% [02:45].
- OSWorld verified (Computer use): Sonnet 5 scores 81.2% vs Sonnet 4.6 at 78.5% and Opus 4.8 at 83.4% [02:45].
- GPQA Diamond (Knowledge work): Sonnet 5 scores 1418 vs Sonnet 4.6 at 1395 and Opus 4.8 at 1615 [02:45].
- Cost vs. Performance: On the BrowseComp benchmark, Sonnet 5 achieves roughly half the cost per task (~$4.50 vs ~$8.00 on medium effort) compared to Opus 4.8 with only about a 5% difference in pass rate [03:19].
- Fable 5 Status: Leaked code strings in Claude Code indicate Fable 5 will require separate credit billing/API usage and US identity verification upon return [10:24].
Notable quotes
- "It is by far the best bang for your buck in AI right now. It has almost the performance of Opus 4.8, but for a fraction of the price." [00:04]
- "When you're doing actual execution, you don't need a ton of compute if the plan mode was done with a lot of compute." [07:44]
- "It is not replacing Opus 4.8 for me. It's only replacing Opus 4.8 for cheap and quick and easy tasks." [11:08]
Assessment
A community review and hands-on workflow tutorial demonstrating practical use cases for Claude Sonnet 5. The Three.js benchmark and Claude Code workflows are shown live in real time on desktop interfaces without misleading edits.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.