I Tested Opus 5.5 vs. GPT-6 Sol on 10 Real Use Cases
Nate Herk | AI Automation · 2026-09-22 · review · 195,343 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Nate Herk from AI Automation Society (AIS) conducts an extensive head-to-head comparison between Anthropic’s Claude Opus 5.5 and OpenAI’s GPT-6 Sol. Across ten complex automation tasks—including web design, video generation, data dashboards, 3D web environments, and browser agents—he tests their output quality, completion speed, and API token costs.
What is shown
- API pricing breakdown [00:16]: Input/output costs per million tokens for Claude Opus 5.5 ($4 input / $20 output) versus GPT-6 Sol ($2 input / $10 output).
- Transcript Search & Ingestion Baseline [01:10]: Both models process 4 hours of meeting transcripts in parallel. Claude Opus 5.5 correctly identifies the latest mention of "n8n" (Sept 14), while GPT-6 Sol misidentifies it as August 17.
- Task 1: Web Design [02:56]: Generating an animated, layered landing page for "Perkform" protein coffee. Opus 5.5 produces realistic 3D bottle rotation and scroll effects; Sol creates flat graphics.
- Task 2: Sizzle Reel Generation [05:02]: Editing 100GB of event footage into a 30-second promotional video using Hyperframes. Opus 5.5 delivers high-energy pacing, b-roll, motion graphics, and audio sync.
- Task 3: Social Video (Reel) [08:02]: Transforming raw video into an edited Instagram Reel explaining Andrej Karpathy's workflow. Opus 5.5 integrates animated UI graphics, captions, and SFX.
- Task 4: Financial Analytics Suite [10:34]: Generating Google Sheets financial models, pitch decks, and KPI dashboards for BrightPath Analytics, revealing that both models overlapped and edited shared workspace files.
- Task 5: 3D Mini-Game [16:03]: Writing a browser-based 3D exploration game ("Small Hours") in Three.js/WebGL with lighting and interactive objects.
- Task 6: Interactive 3D Learning World [18:47]: Synthesizing 100 YouTube video transcripts into a walkable 3D academy with interactive mini-demonstrations of LLM mechanics.
- Task 7: 3D Itinerary Planner [23:00]: Building an interactive 3D globe travel guide covering AI conferences and scenic parks across October.
- Task 8: Codebase Repair Benchmark [26:24]: Evaluating bug-fixing and multi-file code repair capabilities on a large repository. GPT-6 Sol scores 100/100 (30/30 checks passed), beating Opus 5.5 at 96.7/100 (29/30).
- Task 9: Skool Course Upload Browser Agent [28:07]: Controlling browser actions to upload a 15-lesson video curriculum, descriptions, and assets into a Skool community.
- Task 10: Canvas Vector Recreation [30:13]: Using browser tools in Canva to sketch and replicate a reference photo using digital drawing instruments.
Claims & numbers
- The presenter notes Opus 5.5 API pricing is double GPT-6 Sol: $4/$20 per million tokens for Opus versus $2/$10 for Sol [00:26].
- Across the ten test runs, Opus 5.5 won 7 categories, GPT-6 Sol won 1 category (codebase repair), and 2 tasks were deemed ties/invalid due to workspace cross-contamination [31:49].
- Codebase repair benchmark scores: GPT-6 Sol achieved 100/100 and passed 30/30 independent checks in 22m 4s for $1.04; Opus 5.5 scored 96.7/100 passing 29/30 checks in 40m for $19.82 [26:29].
- Cumulative totals across all runs: Claude Opus 5.5 ran for 8 hours, 40 minutes, and 16 seconds, costing $213.03; GPT-6 Sol ran for 5 hours, 51 minutes, and 1 second, costing $74.46 [32:00] (with a noted $19 post-correction on Task 4 [41:10]).
Notable quotes
- "Opus 5.5 is a major step up from Opus 5. GPT-6 Sol is a step down from 5.6 Sol; it feels more like a GPT-6 Luna that might come out." [33:12]
- "I would trust Claude Opus more for creativity and for some judgment calls, and I would maybe want to defer some work to GPT-6 Astra if I know very, very specifically what I want." [33:24]
- "A pretty cool scenario would be using Opus 5.5 as the orchestrator... and sends off very specific instructions to a bunch of little GPT-6 Sol workers." [33:41]
Assessment
This is an authentic, independent empirical benchmark review conducted by an AI workflow practitioner. The video documents real desktop screen captures, terminal logs, code executions, and edge-case execution errors (including local environment collision between parallel agents and browser mouse-capture glitches).
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.