Vibe Coding With Claude Sonnet 5.5
BridgeMind · 2026-09-29 · community · 68,787 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Matthew Miller, founder of BridgeMind, hosts a livestream showcasing and benchmarking AI agent workflows, software development, and the newly released Claude Sonnet 5.5 model. During the broadcast, he tests and compares Sonnet 5.5 against Claude Opus 5.5, Claude Fable 5.1, and GPT-6 Astra across code generation, 3D interactive web applications, Blender model generation, and motion graphics video generation.
What is Shown
- BridgeMind Ecosystem & BridgeVerse [09:15 - 13:30, 71:15 - 73:25]: Demonstrates BridgeMind One's Rust-based client, terminal dashboard, bomb sprint timer, and BridgeVerse—a gamified 3D virtual office space where autonomous AI coding agents (Claude Code, Codex, Grok) sit at desks, work on workspaces, and can be managed via a voice-interactive 3D assistant.
- BridgeBench & Nerf Bench [26:15 - 28:55]: Displays BridgeMind's AI leaderboard tracking model capabilities and Nerf Bench, which tracks post-launch performance degradation (showing Claude Opus 5.5 at 99.2% power and GPT-6 Astra at 102.8% power).
- ElevenLabs v4 Announcement & Testing [108:15 - 110:30, 276:25 - 277:35]: Reviews the release announcement of ElevenLabs' Eleven v4 and v4 Turbo voice models and generates an animated promotional video for BridgeVerse featuring Eleven v4 voiceover and music.
- Claude Sonnet 5.5 Breaking News & Setup [180:30 - 188:55]: Receives live notice of Anthropic dropping Claude Sonnet 5.5 in Claude Code (
v2.1.284), updates the CLI environment, and verifies model availability. - Anthropic Official Benchmarks & Pricing Review [194:15 - 195:35, 226:20 - 228:40]: Examines the official announcement and Artificial Analysis charts showing Sonnet 5.5 outscoring Opus 5.5 on agentic coding (79.4% vs. 66.4%) and matching Fable 5.1 on intelligence index when run at Max effort.
- Design Bench Comparisons on BridgeBench [240:00 - 248:30, 255:05 - 257:30]: Runs and evaluates 3D WebGL/Three.js simulation benchmarks for Sonnet 5.5 against Opus 5.5, Fable 5.1, and GPT-6 Astra:
- Black Hole Merger [240:10]
- Rocket Launch [244:45]
- Lava Lamp [246:40]
- Sunset Ocean [255:10]
- Turntable [255:45]
- Blender 3D Modeling via MCP [249:40 - 251:30]: Inspects a high-detail rocket model generated directly in Blender using a custom Blender MCP tool.
- Playable 3D Games Built with Sonnet 5.5:
- Bridge Horror House [235:40 - 238:40]: A first-person horror survival game generated with medium effort.
- Operation Last Stand / Dead Signal [280:25 - 282:10, 307:30 - 309:05]: A 3D wave-based first-person zombie shooter generated at Max effort ($177 API cost, 49-minute build time), tested live with weapon swapping, sound effects, hit particles, and collision physics.
- Critter Kart Grand Prix [332:40 - 335:05]: A multi-track 3D kart racing game complete with menus, racer selection, AI opponents, sound effects, power-ups, and lap tracking generated in a single shot.
- Code-Generated Motion Graphics Video [341:00 - 344:55]: Displays a 1-minute historical motion graphics video titled "Can machines think? (1943–2026)" built entirely in code by Claude Sonnet 5.5 at Max effort ($25 API cost).
Claims & Numbers
- Productivity & Speed: The presenter claims Claude Opus 5.5 made him approximately 2x to 2.5x more productive in his daily engineering workflows [05:01, 51:10].
- BridgeBench Traffic: The presenter states BridgeBench generated roughly 3 million impressions/views over the preceding week on X [04:00, 61:35].
- Annual Recurring Revenue (ARR): The live stream counter shows BridgeMind's ARR standing at $246,368 to $247,268 during the stream [41:04, 258:20].
- Claude Sonnet 5.5 Benchmarks:
- Anthropic's official performance data shown reports Claude Sonnet 5.5 scoring 79.4% on agentic coding benchmarks at Max effort compared to 66.4% for Claude Opus 5.5 and 54.4% for GPT-6 Astra [194:35, 195:25].
- On Artificial Analysis Intelligence Index, Sonnet 5.5 scores 56 at Max effort (just behind Opus 5.5 at 58) and 52 at Extra High effort [227:15, 232:30].
- On CursorBench 4.0, Sonnet 5.5 Max scores 55.5% ($9.67 cost per task, 271k tokens per task) and Sonnet 5.5 Extra High scores 53.1% ($3.88 cost per task, 100k tokens per task) [200:00, 215:50].
- Pricing & Token Efficiency:
- The presenter notes Sonnet 5.5 costs half the base price of Opus 5.5 ($2/$10 vs. $4/$20 per million input/output tokens) [03:18, 197:30].
- The presenter notes that running Sonnet 5.5 on Max effort is token-intensive, making individual complex tasks cost up to $177 in API credits [311:38, 312:44].
- ElevenLabs v4: The presenter shows ElevenLabs v4 Turbo priced around $0.01 per minute of audio compared to GPT-Live at roughly $0.05 per minute [113:45 - 114:00].
Notable Quotes
- [03:14]: "Sonnet 5.5 is going to be a substantial jump forward, and it's going to jump from F-tier to B-tier. It's going to be priced more than half, or about half the price of Opus 5.5, and it's going to be a very good model."
- [239:00]: "Dude, if that is medium effort... oh my gosh. Okay, we may have a model on our hands. What in the world?"
- [334:25]: "Dude, this is insane! What is going on? ... This is the most complete game that's been created."
Assessment
This is a live, unedited developer stream providing real-time demonstration and benchmarking of AI developer tools and the launch of Claude Sonnet 5.5. The tests, code runs, terminal interactions, and web application executions are conducted live on stream, clearly highlighting both the impressive capabilities of the models (such as single-shot 3D browser games) and their practical tradeoffs, including steep token consumption and multi-minute generation times when using maximum thinking effort.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.