Opus 5.5 vs GPT-6 Sol (Blender F1 Car Test)
Better Stack · 2026-09-29 · review · 11,004 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary A presenter from Better Stack conducts a side-by-side benchmark comparing Claude Opus 5.5, OpenAI GPT-6 Sol, GPT-6 Astra, and Claude Fable 5.1 on 3D Blender modeling and animation tasks. Using identical terminal-based coding agent prompts to research reference photos, construct a detailed Formula 1 car, generate an assembly animation, and animate a pitstop, he evaluates output quality, token usage, cost, and execution time.
What is shown
- [00:17] CLI agent environments: Claude Code running Claude Opus 5.5 (1M context) and OpenAI Codex running GPT-6 Sol, both with extra-high reasoning effort.
- [00:24] The three consecutive prompts: creating a 2026 Ferrari F1 car from web reference images, animating the car assembly, and animating a pit stop sequence.
- [00:34] Terminal logs showing both models browsing the web, downloading SF-26 reference images from Formula 1's website, and noting the user's typo ("F2" instead of "F1").
- [01:17] Blind presentation of "Model 1" results: blueprint-style wireframe assembly animation, high-detail static car renders with carbon fiber texturing and sponsor decals, and pit stop animation with motion blur.
- [02:22] Blind presentation of "Model 2" results: clay/untextured part assembly animation, lower-detail car renders with disconnected parts and inverted decals, and a pit stop animation with floating detached wheels.
- [03:17] Model reveal: Model 1 is Claude Opus 5.5 and Model 2 is GPT-6 Sol.
- [03:28] Token count, price, and runtime breakdown graphics comparing Opus 5.5 and GPT-6 Sol.
- [04:21] Demonstration of GPT-6 Astra: exploded part assembly animation, static render, and an accurate wheel-change pit stop animation.
- [05:06] Demonstration of Claude Fable 5.1: assembly animation, static render showing minor surface artifacts, and a pit stop animation with tire bouncing and chassis suspension.
- [05:47] Four-way split-screen comparison table summarizing renders, costs, and runtimes across all four models.
Claims & numbers
- The presenter states Opus 5.5 and GPT-6 Sol both launched the previous week.
- Both models corrected the prompt's mistaken reference to a "Ferrari 2026 F2 car" by identifying the SF-26 Formula 1 car [00:34].
- Claude Opus 5.5 run stats: 53.6M input tokens (52.2M cached reads), 329K output tokens, $27.86 API cost (including $10.83 in cache writes), and 1 hour 7 minutes active work time [03:29, 03:54].
- GPT-6 Sol run stats: 15.3M input tokens (14.8M cached reads), 57K output tokens, $4.50 API cost, and 48 minutes 44 seconds active work time [03:38].
- Per-token pricing cited: GPT-6 Sol is $2 / $10 (per million input/output tokens), while Opus 5.5 is $4 / $20 [03:46].
- GPT-6 Astra run stats: 22.9M input tokens (22.5M cached reads), 122K output tokens, $32.47 API cost, and 1 hour 49 minutes active work time [05:38].
- Claude Fable 5.1 run stats: 24.9M input tokens (23.7M cached reads), 262K output tokens, $42.44 API cost, and 1 hour 9 minutes active work time [05:44].
- An internal staff poll and YouTube community poll both ranked GPT-6 Astra's pit stop animation first, with Opus 5.5 finishing in a close second place [04:46].
Notable quotes
- [00:08] "Spoiler alert, one of these new models absolutely dominates the other."
- [01:31] "I must say, this is one of, if not the best render I have ever had a model make."
- [06:11] "Opus 5.5 is my new daily driver, and I've not found the need to use Fable while using it."
Assessment This is an authentic third-party benchmark and comparative review demonstrating autonomous coding agents using Python to script Blender 3D assets and animations. The presenter provides clear proof of agent terminal interactions, reproducible prompts, and granular API billing and execution metrics.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.