Claude Opus 5.5 vs GPT-6 Astra: Same 3D Prompt, We Played Both
Lite AI Lab · 2026-09-24 · review · 18,222 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
In this hands-on comparison by Lite AI Lab, Anthropic’s Claude Opus 5.5 and OpenAI’s GPT-6 Astra compete head-to-head in a one-shot coding challenge using the OpenCode agent. Both models are given identical prompts to generate an interactive 3D underwater coral reef and a playable beach buggy racing game in Three.js, testing coding quality, visual aesthetic, cost, thinking tokens, and actual gameplay feel.
What is shown
- [00:20] Pricing and model comparison on OpenRouter: GPT-6 Astra ($10/$50 per 1M tokens) vs. Claude Opus 5.5 ($4/$20 per 1M tokens).
- [00:40] Configuration of OpenCode (version 1.18.32), setting both to default reasoning effort via OpenRouter.
- [01:23] Round 1: Underwater coral reef prompt requiring rocks, arches, 5+ coral species, marine life, and a fish-scattering click interaction.
- [01:58] Demonstration of OpenCode's default 32,000-token output limit causing Opus to stop mid-thinking, and the PowerShell environment variable fix (
$env:OPENCODE_EXPERIMENTAL_OUTPUT_TOKEN_MAX = "128000"). - [02:58] Autonomous browser inspection loop: both models open their generated HTML in a headless/automated browser, take screenshots, review visual layouts, and iterate.
- [03:22] Side-by-side visual evaluation of the generated coral reefs: Opus produces vibrant colors and 150 fish, while Astra creates a documentary-style design named "Pelagic / The hidden lagoon" with interactive UI and 64 fish.
- [04:10] The click test: testing fish scattering and regrouping behavior in both reef simulations.
- [05:13] Round 2: 3D beach buggy racing game prompt requiring jump ramps, 3 AI opponents, track obstacles, power-ups, HUD, and restart mechanics.
- [06:11] Live playthrough of Claude Opus 5.5’s game ("Tropical Buggy Rush"), showcasing custom engine sounds, shockwaves, shield power-ups, and vehicle physics.
- [06:40] Live playthrough of GPT-6 Astra’s game ("Dune Breaker"), showcasing a stylized retro start screen, jump airtime tracking, and power-up boosts.
- [07:30] Artificial Analysis benchmark data and overall token/cost summary across both coding tasks.
Claims & numbers
- The presenter says Claude Opus 5.5 token pricing is $4/1M input and $20/1M output, making its base token rate 60% cheaper than GPT-6 Astra at $10/1M input and $50/1M output [00:20].
- The presenter notes OpenCode caps output tokens per response at 32,000 tokens by default, which Opus exceeded during reasoning [02:05].
- For Round 1 (Coral Reef), the presenter states GPT-6 Astra finished in ~16 minutes costing $3.53, while Claude Opus 5.5 took 25.5 minutes costing $3.46 [03:08].
- For Round 2 (Racing Game), Astra completed in ~20 minutes costing $5.06 (82.3 KB, 524 lines of code), while Opus took ~29 minutes costing $5.75 (110.8 KB, 1,078 lines of code) [05:53].
- Citing Artificial Analysis Intelligence Index benchmarks (checked 2026-09-23), the presenter reports Claude Opus 5.5 is ranked #1 with a score of 58 (out of 210 models), while GPT-6 Astra ranks #6 with a score of 53 [07:34].
- Across the two tasks, Opus consumed 179,282 thinking tokens, whereas Astra used only 17,568 thinking tokens (roughly 10 times fewer) [07:58].
- The presenter highlights that despite cheaper per-token rates, Opus was more expensive overall due to extensive thinking tokens: Opus cost $9.21 (57m 18s active time) vs. Astra's $8.59 (35m 54s active time) [08:08].
Notable quotes
- [00:00] "One prompt, two AIs, and no second chances. Each one gets a single try to build a 3D world."
- [00:34] "A cheaper token doesn't always mean a cheaper bill. Which one is actually worth your money?"
- [07:14] "Astra's game looks a little better, but Opus's game plays better. The driving feels more natural, and the sound is better too."
Assessment
This is an independent, detailed third-party review and direct head-to-head benchmark. The creator transparently presents real-time terminal output, unmodified code execution, browser console error logs (both scoring zero errors), and hands-on gameplay mechanics to validate performance.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.