I Made Claude Opus 5.5 & GPT 6 Astra Build the Same App (Raw Results)
Dubibubi · 2026-09-23 · community · 74,450 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Dubibubi conducts a head-to-head evaluation comparing Anthropic's Claude Opus 5.5 and OpenAI's frontier model GPT-6 Astra, running both on maximum effort. The models compete across three tasks: building a competitor intelligence web application, coding a stop-motion animated short within a single HTML file, and performing automated code review with cross-verification.
What is shown
- [00:15] Overview of the competitive context, showing OpenAI's release of GPT-6 Sol and Luna shortly after the Claude Opus 5.5 launch, referencing Terminal-Bench 4.0 scores.
- [01:47] Test setup running Claude Opus 5.5 in Claude Code against GPT-6 Astra in Codex, both configured to maximum effort mode.
- [04:42] Evaluation of Opus 5.5’s competitor intelligence app ("Signal"), featuring video breakdown, hook analysis, viral script suggestions, and metrics extraction.
- [08:36] Review of GPT-6 Astra's competitor app, highlighting a cleaner, less cluttered interface but comparatively shallower analytical breakdowns.
- [11:31] First build metrics and scoreboard: Opus 5.5 used 118M tokens ($40.20 API cost, 1h 24m) vs. Astra's 29M tokens ($43.61 API cost, 1h 27m).
- [13:25] Test 2 creative coding prompt ("Keep Zooming") asking both models to generate a continuous zoom stop-motion animation using only HTML, CSS, and JavaScript.
- [15:08] Playback of GPT-6 Astra's generated stop-motion animation, displaying basic vector graphics and minor anatomical visual flaws.
- [16:16] Playback of Opus 5.5's stop-motion animation, demonstrating fluid multi-scale zooming, textured illustrations, and integrated sound design.
- [17:02] Test 2 metrics: Astra finished in 25m 27s for $8.15, while Opus 5.5 took 1h 37m and cost $23.33.
- [18:16] Test 3 multi-agent code review setup, where each model audits a codebase and verifies the other's reported bugs (+1 for confirmed bug, -1 for hallucinated bug).
- [19:41] Verification results: Astra confirms 10 of 10 bugs submitted by Opus 5.5; Opus 5.5 confirms 23 of 23 bugs submitted by Astra.
- [21:46] Final scoreboard reveal showing a 6–6 deadlock tie across the three rounds.
Claims & numbers
- The presenter notes Terminal-Bench 4.0 benchmarks placed Opus 5.5 significantly ahead of GPT-6 Astra at extra-high and maximum effort settings.
- In Test 1, Opus 5.5 consumed 118,536,498 total tokens ($40.20 API equivalent cost) across 1 hour 24 minutes, while GPT-6 Astra used 29,203,456 tokens ($43.61 cost) across 1 hour 27 minutes.
- In Test 2, Astra completed the animation in 25 minutes 27 seconds for $8.15 (4,402,824 tokens), saving 65.1% in cost compared to Opus 5.5, which required 1 hour 37 minutes and cost $23.33 (36,419,787 tokens).
- In Test 3, GPT-6 Astra identified 23 verified bugs at a cost of $17.66 (10,668,804 tokens, 34m 16s), whereas Opus 5.5 found 10 verified bugs costing $29.24 (90,403,832 tokens, 25m 38s).
- Overall costs across all tests: Opus 5.5 totaled $92.77 in estimated API cost across 5h 55m of compute, while GPT-6 Astra totaled $69.42 across 2h 26m.
Notable quotes
- [07:05] "This is such a finished product. I feel like I could literally charge like $10 a month for this."
- [20:07] "Dude, Astra had 23 confirmed bugs. That means its final score is 23 points."
- [21:55] "We have a deadlock tie. Now, I promise you this is not scripted, but honestly, you can really see the different strengths here."
Assessment
This is an independent hands-on technical benchmark and comparison video by a software developer testing Claude Opus 5.5 and GPT-6 Astra side by side. All apps, animations, and token logs are demonstrated on-screen in real time or full screen without deceptive edits, accurately capturing each model's distinct tradeoffs in creative output quality versus token efficiency.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.