I Tested Opus 5.5 vs GPT-6 Astra (CLEAR Winner)
Jack Roberts · 2026-09-24 · review · 35,631 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
In this video, creator Jack Roberts compares Anthropic’s Claude Opus 5.5 against OpenAI’s GPT-6 Astra across five real-world coding, animation, and design tasks. Using identical prompts and a $100 budget per model, he tests both systems on web design, launch video recreation, pure JavaScript animation, a browser ninja game, and brand identity design.
What is shown
- Benchmark overview [00:23]: Presentation slides detailing performance, Terminal-Bench 4.0 accuracy vs. cost, and OpenAI pricing charts comparing GPT-6 Astra, GPT-6 Sol, and GPT-6 Luna.
- Task 1: Website from scratch [01:23]: Jack compares full personal website redesigns generated by Astra and Opus 5.5 using video and image assets generated via the Higgsfield API.
- Task 2: Remake launch film [04:30]: A five-second brand film recreation prompt given to both models; Jack reviews the visual timing and integrated text graphics [05:18].
- Task 3: Glaido film in pure code [06:17]: Both models generate an animated promotional short purely in JavaScript code without video generators. Astra outputs a 2D floating ghost animation [06:42], while Opus 5.5 produces an animated cartoon character ("Pip") with music, sound effects, typing effects, and UI transitions [07:15].
- Task 4: Playable ninja game [08:50]: Both models create a playable 2D browser platformer game ("Moonblade"). Astra's version features jumping and guard-clearing mechanics [09:00], while Opus 5.5 includes double jumping, archers, slice animations, sound effects, and combat pacing [09:21].
- Task 5: Brand identity board [10:14]: Evaluating brand design boards for "Stacked AI", inspecting color palettes, logo mockups, and typography layouts [10:40].
- Course and Agentic OS overview [10:47]: Brief walkthrough of Jack's "Claude Code Full Course" and his custom "Agentic OS" multi-model workflow setup.
Claims & numbers
- The presenter states that Claude Opus 5.5 is 40% cheaper and roughly 30% faster than Claude Fable 5.1.
- A benchmark slide states Opus 5.5 matches GPT-6 Astra on Terminal-Bench 4.0 for about 40% of the cost.
- Pricing shown for OpenAI's GPT-6 lineup: Astra at $50 per 1M tokens, Sol at $10 per 1M tokens (5x cheaper), and Luna at $0.50 per 1M tokens (100x cheaper).
- The presenter runs the comparison across five real tasks under identical prompts and a $100 credit budget.
- Scoring outcome: Opus 5.5 wins Website (Task 1), Glaido Film (Task 3), and Ninja Game (Task 4); Launch Film (Task 2) and Brand Board (Task 5) are ruled ties, concluding in a 3–0 win for Opus 5.5.
Notable quotes
- [00:00] "Opus 5.5 is 40% cheaper than Fable and 30% faster, and in this video, we're going to compare it against Astra to see which model is better."
- [08:18] "That is a clear and unequivocal win for Opus 5.5. That has actually genuinely blown me away. That is a new capability."
- [14:38] "A combination of both Astra and Opus is exactly where you want to be."
Assessment
This is a hands-on independent review and comparison video featuring side-by-side execution of real prompts in browser environments. All five coding and design deliverables (websites, JavaScript animations, and playable canvas games) are demonstrated running directly on screen, with straightforward, subjective judging by the host.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.