Claude Sonnet 5.5 Just Beat Opus at Coding... Then It Built All This
Tech2WiLD · 2026-09-29 · community · 3,936 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
In this review and demo video, tech creator Tony (Tech2WiLD) discusses Anthropic’s release of Claude Sonnet 5.5, analyzing its benchmark results, cost-efficiency curves, and leaderboard positions relative to Claude Opus 5.5 and OpenAI's GPT-6 Sol. He also demonstrates four distinct interactive applications built using Claude Sonnet 5.5, ranging from a political news aggregator to complex 3D voxel simulators and games.
What is shown
- [00:00] Intro & Context: The presenter introduces Claude Sonnet 5.5 following its release announcement on Anthropic's blog.
- [01:13] Benchmarks Table Review: Walkthrough of official performance figures comparing Sonnet 5.5 against Sonnet 5, Opus 5.5, and GPT-6 Sol across agentic coding (Terminal-Bench 4.0, FrontierCode, Cursor-Bench), knowledge work, and reasoning.
- [03:14] Cost vs. Performance Curves: Examination of Terminal-Bench 4.0 score vs. cost per task across effort levels (Low, Medium, High, XHigh, Max).
- [04:22] Self-Edited Video: The presenter discloses that the cuts and editing of the video itself were automated by Claude Sonnet 5.5 using his local editing engine.
- [04:56] Artificial Analysis Leaderboard: Review of the Artificial Analysis Intelligence Index, showing Sonnet 5.5 jumping from index 38 to 56, landing at rank #2 behind Opus 5.5 (58).
- [08:18 - 09:42] Cost per Task Breakdown: Analysis of cost per task across effort tiers on Artificial Analysis, contrasting Sonnet 5.5 and Opus 5.5 at Max vs. High/Medium settings.
- [10:02] Demo 1: News Pipeline: A live browser-based news aggregation application classifying 608 articles into Conservative (156), Independent (147), and Progressive (141) ideological perspectives with clustering and deduplication.
- [12:01] Demo 2: Six Flags Over Georgia Voxel Simulation: A full 3D interactive voxel recreation of the theme park in Mableton, Georgia, featuring a playable first-person roller coaster ride on Goliath and flyover mode.
- [14:50] Demo 3: Sky City: A 3D voxel city simulation where autonomous NPCs converse via locally hosted models (Qwen3.8-27B and MiniMax M3 running on dual RTX 3090 GPUs), with interactive sandbox tools including a destructive tornado and UFO tractor-beam abductions.
- [17:18] Demo 4: Warzone Voxel: A detailed destructible 3D military simulation with coastline, battleships, helicopters, airstrikes, and callable Tomahawk missile strikes.
Claims & numbers
- Release Timing: The presenter states Claude Sonnet 5.5 dropped on September 28, 2026, roughly three months after Claude Sonnet 5 (released June 30, 2026).
- Speed & Pricing: The presenter highlights that Sonnet 5.5 runs 30%+ faster than Sonnet 5 and costs up to 30% less for most tasks ($2.00 / 1M input tokens, $10.00 / 1M output tokens, $0.20 cache read, $2.50 cache write).
- Terminal-Bench 4.0 Coding: Sonnet 5.5 scored 70.6%, outperforming Claude Opus 5.5 (66.8%), GPT-6 Sol (58.0%), and Sonnet 5 (10.3%).
- FrontierCode (0-shot): Sonnet 5.5 achieved 46.2%, compared to 42.4% on Sonnet 5, 54.4% on Opus 5.5, and 64.9% on GPT-6 Sol.
- Cursor-Bench: Sonnet 5.5 achieved 55.0% vs. 41.5% for Sonnet 5, 57.8% for Opus 5.5, and 58.2% for GPT-6 Sol.
- Artificial Analysis Intelligence Index: Sonnet 5.5 reached an index score of 56 (up 18 points from Sonnet 5's 38), ranking #2 overall right behind Opus 5.5 (58).
- Cost Discrepancy at Max Effort: The presenter notes on Artificial Analysis that at "Max effort," Sonnet 5.5 becomes substantially more expensive ($14.60 per task) than Opus 5.5 at Max ($8.86 per task) due to high verbosity and output token counts, but at "High effort" Sonnet 5.5 costs only $1.00 to $2.74 compared to Opus 5.5's higher base cost.
- Context Window: Sonnet 5.5 features a 1 million token context window.
Notable quotes
- [00:00] "Claude just dropped Sonnet 5.5, and they are basically saying that they're taking the lead when it comes to the frontier."
- [04:22] "Well, this video you see right now was fully edited by Sonnet 5.5."
- [18:47] "It seemed like Dario saw something within these his new models, internal models, and we're seeing now... it obviously seems like he got something going on that is just going to outpace the competition."
Assessment
This is a genuine third-party review and independent benchmark/demo showcase by creator Tech2WiLD. The benchmark charts and pricing tables reflect Anthropic and Artificial Analysis evaluations, while the four live-running voxel web apps and news aggregation tool provide functional, unedited demonstrations of Sonnet 5.5's coding outputs.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.