I Made Opus 5.5, Fable 5.1 & GPT-6 Build the Same App (RAW RESULTS)
Pat Simmons · 2026-09-25 · community · 143,546 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Pat Simmons conducts a head-to-head evaluation comparing three frontier AI models—Claude Opus 5.5, Claude Fable 5.1, and GPT-6 Astra—on three complex coding and creative tasks under a strict "one prompt, zero human revisions" protocol. Across three builds (a procedural risograph storybook, an animated macro launch film and landing page, and a playable 3D Tony Hawk’s Pro Skater clone), Simmons inspects the generated output quality, execution times, and calculated API token costs.
What is shown
- 00:00 – 00:43: Introduction of the three models and benchmark parameters: Claude Fable 5.1 (left), GPT-6 Astra (middle), and Claude Opus 5.5 (right), all running in agentic CLI harnesses at high effort level.
- 00:44 – 02:15: Build 1 Prompt Setup: A procedural Moby-Dick scene explorer rendered in a simulated risograph print style as a single-file HTML/JS canvas app without external image generation, inspired by Kevin Ngo's 25-room Opus 5 experiment.
- 02:23 – 09:30: Build 1 Results:
- [03:07] GPT-6 Astra’s generated Moby-Dick interactive diorama.
- [05:06] Claude Fable 5.1’s version with walking animation and chapter popups.
- [06:51] Claude Opus 5.5’s version featuring intricate isometric scenes (New Bedford, the chapel, the Spouter-Inn, the Pequod deck, animated swimming whales).
- [09:17] Session log cost analysis for Build 1.
- 10:29 – 11:18: Build 2 Prompt Setup: Recreation of Anthropic's Claude Opus 5.5 "microscopic horizon" launch video and announcement webpage, generating imagery via GPT Image and synthesizing all sound effects in code.
- 11:19 – 20:45: Build 2 Results:
- [11:19] Fable 5.1's version ("Loupe"), showing macro images and abrasive sound design.
- [13:12] Astra's version ("Loam"), featuring moss and fungi imagery with subtle audio.
- [15:52] Opus 5.5's version ("Terra Minima" / Halden Optical), featuring curved horizons, matched-cut rotating frames, procedural synth audio, and an accurate website layout.
- [20:46] Session log cost analysis for Build 2.
- 21:09 – 25:15: Build 3 Prompt Setup: Creating a 3D Tony Hawk's Pro Skater warehouse level clone using headless Blender via Python scripts to model/rig an anatomically proportioned skater and warehouse, exported to GLB and loaded into a playable Three.js web game.
- 25:18 – 34:10: Build 3 Results & Gameplay:
- [25:19] Astra's game ("Opening the Warehouse"), demonstrating functional skating, kickflips, and bails.
- [28:19] Fable 5.1's game ("Warehouse Pro Skater"), showing higher texture fidelity and jumping physics, despite visual glitches with skater hands.
- [30:41] Opus 5.5's game ("Late Shift: Warehouse Session"), featuring volumetric lighting, realistic skater geometry, rail grinding balance meter, drop-ins, and THPS-accurate physics.
- 34:11 – 36:02: Final cost breakdown, summary of model strengths, and closing remarks.
Claims & numbers
- Build 1 (Moby-Dick Risograph):
- GPT-6 Astra finished in 26 minutes (73,581-byte HTML file), generating 16 animated scenes; calculated API cost was $6.20 (or $8.16 including deployment tokens).
- Claude Fable 5.1 finished in ~1 hour; calculated API cost was $42.16.
- Claude Opus 5.5 finished in ~1 hour 10 minutes (after a 30-minute usage limit reset wait); calculated API cost was $25.49.
- Build 2 (Launch Film & Site):
- GPT-6 Astra finished in 26 minutes; calculated API cost was $7.96 ($47.37 without prompt caching).
- Claude Fable 5.1 finished in 26 minutes; calculated API cost was $16.33.
- Claude Opus 5.5 finished in ~40 minutes; calculated API cost was $11.66 ($69.83 without prompt caching).
- Build 3 (Tony Hawk's Pro Skater 3D Game):
- GPT-6 Astra completed initial gameplay in 10 minutes and full build in 48 minutes; calculated API cost was $44.08.
- Claude Fable 5.1 finished in 56 minutes; calculated API cost was $32.49.
- Claude Opus 5.5 finished in approximately 2 hours; calculated API cost was $58.05.
- The presenter notes he is testing using 20x subscription tiers for both ChatGPT and Claude.
Notable quotes
- [08:09]: "Geez, okay, Opus clearly won that one... just, without a doubt, winner there."
- [34:12]: "So there we go: Opus 5.5 across the board seems to be the clear winner."
- [34:44]: "And to be clear too, I'm still partial to Astra in my day-to-day... I really like how methodical Astra is. Rarely do I have to come back and say, you know, 'you did this wrong' or have any kind of feedback."
Assessment
This is a real, hands-on independent review and technical demonstration by a community developer running live autonomous software agents across frontier models. The video records full browser interactions and gameplay directly from terminal agent outputs without apparent deceptive staging or skipped runtime discrepancies.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.