I Tested Sonnet 5.5 vs Opus 5.5 vs GPT 6 Astra (No Hype Assessment)
Chase AI · 2026-09-29 · review · 29,835 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Chase Hanninger (Chase AI) conducts a hands-on head-to-head comparison of three frontier AI models—Claude Sonnet 5.5, Claude Opus 5.5, and OpenAI's GPT-6 Astra. Through four practical web development and coding tests (JavaScript animation, landing page UI design, interactive 3D globe visualization, and a Three.js tank game), he evaluates their real-world capabilities, aesthetics, and token costs to determine the best model for developers.
What is shown
- Benchmark & Pricing Overview [00:47]: Comparison table reviewing published benchmarks (Terminal-Bench 4.0, FrontierCode 1.1, Humanity's Last Exam, GDPval-AA, AA Index) and API pricing for Sonnet 5.5 ($2/$10), Opus 5.5 ($4/$20), and GPT-6 Astra ($10/$50).
- Test 1: Pure JavaScript 15-Second Explainer Animation [01:32]:
- Prompt asking models to code an animated explainer in JavaScript showing how Claude subagents preserve context memory.
- Sonnet 5.5 output demonstration [02:08].
- Opus 5.5 output demonstration [02:49].
- GPT-6 Astra output demonstration [03:19].
- Test 2: Boutique Hotel ("Dune House") Landing Page [04:14]:
- Use of the Higgsfield API/MCP for image generation alongside Anthropic models [04:29].
- Sonnet 5.5 landing page layout with full-width hero header [05:02].
- Opus 5.5 landing page featuring interactive mouse-over effects, custom logo, glassmorphism, and room selection [06:14].
- GPT-6 Astra landing page with clean hero imagery and card layouts [07:51].
- Sponsorship / Chase AI+ Demo [09:09]: Showcase of the Chase AI+ classroom, Claude Code and Codex masterclasses, and the "Jarvis" agentic OS interface.
- Test 3: 3D Sci-Fi Interactive Travel Dashboard ("Meridian") [09:33]:
- Sonnet 5.5 generating "Meridian" with flight paths, interactive zoom, and destination city views [09:59].
- Opus 5.5 generating "Meridian Flight Atlas" with a flat polar view toggle and city inspection cards [11:09].
- GPT-6 Astra generating "Orbit", a functional travel booking dashboard with practical trip-planning controls [12:22].
- Test 4: Browser-Based 3D Tank Game in Three.js [13:38]:
- Sonnet 5.5's "Iron Vanguard", testing garage tank selection, projectile ballistics, sniper zoom, and bot battle [13:49].
- Opus 5.5's "Steel Vanguard", testing tank armor stats, vehicle driving, and destructible elements [14:54].
- GPT-6 Astra's "Iron Meridian", featuring tactical battle maps, ricochet angle physics, and bot encounters [15:56].
Claims & numbers
- The presenter displays published benchmark scores [00:47]:
- Terminal-Bench 4.0: Sonnet 5.5 scored 70.6%, Opus 5.5 scored 66.4%, GPT-6 Astra scored 57.9%.
- FrontierCode 1.1 (Main): Opus 5.5 scored 54.4%, GPT-6 Astra scored 53.3%, Sonnet 5.5 scored 46.2%.
- Humanity's Last Exam: Opus 5.5 scored 67.7%, Sonnet 5.5 scored 64.5%, GPT-6 Astra scored 57.2%.
- GDPval-AA: Opus 5.5 scored 1846, Sonnet 5.5 scored 1844, GPT-6 Astra scored 1542.
- Artificial Analysis Index: Opus 5.5 scored 58, Sonnet 5.5 scored 56, GPT-6 Astra scored 53.
- The presenter states model API pricing per million tokens [01:11]:
- Sonnet 5.5: $2 input / $10 output.
- Opus 5.5: $4 input / $20 output (double Sonnet 5.5).
- GPT-6 Astra: $10 input / $50 output (five times Sonnet 5.5).
- The presenter notes token consumption per test:
- Test 1: Opus 5.5 and Sonnet 5.5 used ~300,000 tokens; GPT-6 Astra used ~70,000 tokens [04:02].
- Test 2: Opus 5.5 and Sonnet 5.5 used ~200,000 tokens; GPT-6 Astra used ~125,000 tokens [08:58].
- Test 3: Opus 5.5 and Sonnet 5.5 used ~300,000 tokens; GPT-6 Astra used ~150,000 tokens [13:31].
- Test 4: Sonnet 5.5 used ~750,000 tokens; Opus 5.5 used ~500,000 tokens; GPT-6 Astra used ~300,000 tokens [14:58, 16:17].
Notable quotes
- [00:26] "In fact, when we look at something like Sonnet 5.5, it actually posts better benchmarks at agentic coding than its bigger brother, Opus."
- [03:36] "GPT-6 definitely leaves something to be desired when we compare this to both Opus and Sonnet—not nearly as dynamic."
- [17:28] "Overall, when we take all these benchmarks into account, I think the winner here is Opus 5.5, but the other two models, Astra and Sonnet, are not far behind."
Assessment
This is an authentic, independent technical review demonstrating real browser applications and scripts generated by Claude Sonnet 5.5, Claude Opus 5.5, and GPT-6 Astra. The creator shows live, functional software execution in the browser across all four prompts, candidly reporting token usage and qualitative differences without unsubstantiated claims.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.