I Tested Opus 5.5 vs. GPT-6 Astra on 12 Real Use Cases
Nate Herk | AI Automation · 2026-09-24 · review · 124,876 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
In this video, creator Nate Herk conducts an extensive head-to-head benchmark comparing Anthropic's Claude Opus 5.5 and OpenAI's GPT-6 Astra across 12 real-world use cases. Testing tasks ranging from website generation and video editing to 3D world creation and complex codebase refactoring, Herk evaluates each model's speed, API-equivalent cost, and qualitative output.
What is shown
- Cost & Setup Overview [00:33]: API billing comparison ($4 input / $20 output per million tokens for Opus 5.5 vs. $10 input / $50 output per million tokens for Astra) running on "High" effort settings.
- Test 1: Perkform Coffee Landing Page [01:39]: Opus creates a dark-themed interactive landing page with 3D product animations (40m 21s, $18.32); Astra creates a light, clean alternative with interactive flavor selectors (32m 23s, $11.33). Opus wins on visual design.
- Test 2: Event Sizzle Reel [04:36]: Editing 105 GB of conference footage into a 30-second promo via Hyperframes. Opus (31m 32s, $10.36) beats Astra (39m 27s, $21.85) in rhythmic pacing and motion layering.
- Test 3: Explainer Reel [07:25]: Generating an Instagram reel summarizing Andrej Karpathy's 3-layer system. Opus (40m 14s, $11.14) produces dynamic motion graphics, beating Astra's simpler edit (22m 38s, $8.91).
- Test 4: BrightPath Analytics Multi-Deliverable [10:19]: Building a 17-slide pitch deck, multi-tab financial model in Google Sheets, dashboard, and landing page. Opus (39m 20s, $17.50) edges out Astra (46m 15s, $16.61) due to richer formulas and narrative depth.
- Test 5: 3D Miniature Museum Escape Game [17:48]: Opus (1h 32m, $31.27) generates a full first-person 3D flashlight escape room; Astra (34m 34s, $7.92) creates an isometric point-and-click puzzle game. Astra wins on execution speed and cost efficiency.
- Test 6: 3D Educational Campus [22:14]: Processing 100 YouTube video transcripts into interactive 3D learning worlds ("Curiosity Campus" vs. "AI Explorer Academy"). Opus (1h 44m, $60.53) wins on depth over Astra (45m 00s, $12.43).
- Test 7: 3D Interactive Travel Itinerary [27:23]: Building a month-long trip planner with an interactive globe and direct flight/hotel booking links. Astra's "Atlas" (32m 07s, $10.99) wins over Opus's "October Journey" (29m 11s, $14.47).
- Test 8: Synthetic Codebase Challenge [30:06]: A test suite designed by Grok and audited by Claude Fable 5.1 and GPT-6 Sol. Astra (35m 13s, $9.14) completes it dramatically faster than Opus (2h 29m, $17.48), winning the round.
- Test 9: Animated Biography Reel [31:44]: Generating a 30-second animated story of Nate Herk. Opus creates a 3D Pixar-style render with voice cloning (48m 23s, $7.76), winning over Astra's claymation-style reel (19m 38s, $11.14).
- Test 10: Browser Canvas Drawing Recreation [34:40]: Recreating a photograph of Nate Herk with Adam Sandler inside Canva using drawing tools. Opus (41m 24s, $8.65) achieves a recognizable likeness, while Astra (26m 51s, $9.96) produces a distorted output.
- Test 11: Social Carousel [36:28]: Formatting a Polymarket polling tweet into an educational slide carousel. Astra (12m 48s, $6.82) wins over Opus (23m 37s, $10.42).
- Test 12: Book Sales Page [38:22]: Redesigning a book landing page for Becoming AI Native. Opus (14m 55s, $6.64) wins for richer storytelling over Astra (14m 11s, $5.32).
- Overall Metrics & Tally [40:38]: Claude Opus 5.5 wins 8–4 against GPT-6 Astra. Astra is 44.8% faster in total runtime (6h 01m vs. 10h 53m) and 38.3% cheaper ($132.43 vs. $214.54).
Claims & numbers
- The presenter states that on API pricing, Claude Opus 5.5 costs $4/million input tokens and $20/million output tokens, while GPT-6 Astra costs $10/million input tokens and $50/million output tokens (2.5× higher token pricing) [00:33, 01:00].
- The presenter reports that across all 12 benchmarks combined:
- Opus 5.5 had a total active runtime of 10 hours, 53 minutes, and 57 seconds [40:50].
- GPT-6 Astra had a total active runtime of 6 hours, 1 minute, and 5 seconds (saving 4 hours, 52 minutes, 52 seconds, or 44.8% less time) [40:50].
- Opus 5.5 total API-equivalent cost was $214.54 [40:50].
- GPT-6 Astra total API-equivalent cost was $132.43 (saving $82.11, or 38.3% cheaper) [40:50].
- In the codebase evaluation (Test 8), the presenter reports that both models passed all 50 independent predetermined tests with a 100/100 score, though an evaluation agent deducted two points from Astra for larger structured test depth [30:35, 30:43].
Notable quotes
- [01:00] "What's really interesting is that Astra is 2.5 times more expensive than Opus 5.5. So, is it going to perform 2.5 times better than Opus 5.5? That's what we're going to see."
- [07:00] "In general, it feels to me like Opus and Claude models are just way more creative and have better, I don't know, taste in a lot of ways..."
- [41:36] "...Opus and Claude models feel like a wise old owl. They feel like they have good judgment and creativity and taste, and GPT models just feel like they are a really good obedient worker."
Assessment
This is an authentic, hands-on practitioner benchmark review comparing frontier models within coding and agentic environments. The presenter demonstrates fully functional live browser apps, scripts, and rendered media while transparently recording runtime lengths and calculated API costs.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.