Opus 5.5 vs GPT-6 Astra build the same physics Rube Goldberg machine: cost and time compared
@leogao25X1,420 views as of 10 October 2026
Why it is here
Curated in github.com/athemeroy/awesome-claude-5-5-videos (‘Mixed workflows and comparisons’). Post claims Opus 5.5 cost $14.28 and 62 min vs GPT-6 Astra $2.20 for the same prompt, despite lower per-token price (single user’s test, unverified).
Description
Description written by Gemini from the videoGemini 3.8 Flash, 10 October 2026
Summary
This side-by-side comparison video, posted by user @leogao25 on September 22, 2026, compares Anthropic’s Claude Opus 5.5 and OpenAI’s GPT-6 Astra coding and executing a self-running 3D physics Rube Goldberg machine from scratch. Both models were given a single high-effort run from an empty directory using their respective developer harnesses (Claude Code vs. Codex CLI), with the screen showing the resulting 3D simulations alongside performance, token, and cost metrics.
What is shown
- Prompt & benchmark setup: Text at the top states: “Opus 5.5 vs GPT-6 Astra: same prompt, same machine” with the prompt: “Build a self-running Rube Goldberg machine. Physics from scratch. One attempt per model, high effort.” [00:00].
- Left pane (Claude Opus 5.5 via Claude Code):
- [00:00–00:06]: A metal sphere drops down a multi-level wooden ramp within a detailed wooden room/cabinet enclosure.
- [00:07–00:10]: The ball rolls across a wooden tabletop, striking and toppling a line of dominoes.
- [00:11–00:16]: The falling domino triggers a counterweight pulley system that raises a checkered finish flag up a flagpole.
- Right pane (GPT-6 Astra via Codex CLI):
- [00:00–00:06]: A marble rolls through elevated green curved wireframe tracks on a minimalist green plane.
- [00:07–00:10]: The marble trips a row of standing upright pegs and a seesaw rocker lever.
- [00:11–00:16]: The kinetic impulse triggers a pulley contraption that hoists an orange triangular target flag.
- Comparative performance table: Detailed breakdown showing build time, API cost, output token count, and pricing tiers [00:00–00:16].
Claims & numbers
- Claude Opus 5.5:
- Harness: Claude Code.
- Build time: 62 minutes.
- API cost: $14.28.
- Output tokens: 265K.
- List price per million tokens: $4 input / $20 output.
- GPT-6 Astra:
- Harness: Codex CLI.
- Build time: 11 minutes.
- API cost: $2.20.
- Output tokens: 20K.
- List price per million tokens: $10 input / $50 output.
- Ratio (Opus 5.5 relative to Astra):
- Build time: 5.8× longer.
- API cost: 6.5× higher.
- Output tokens: 13.6× more.
- List price ratio: 0.4× (cheaper list price per token).
- Run conditions: Stated at the bottom: “Each model started in an empty folder with no config files. Cost is API list price for every token used, including the agent’s pre-training. Time is wall-clock from prompt to finished file. One run each.”
Notable quotes
- “Build a self-running Rube Goldberg machine. Physics from scratch. One attempt per model, high effort.” [00:00]
- “Each model started in an empty folder with no config files.” [00:00]
- “Time is wall-clock from prompt to finished file. One run each.” [00:00]
Assessment
This is a real head-to-head developer benchmark of two frontier models building complex interactive simulations autonomously. While both successfully produced functioning physics contraptions from a cold start, the test reflects a single trial (One run each) and compares two distinct agent harnesses (Claude Code vs. Codex CLI), which significantly influences token volume, reasoning loops, and wall-clock time alongside raw model capability.
Lyrics & themes
- The video is completely silent with no audio, narration, or soundtrack.
- Key themes focus on real-world coding agent cost efficiency, speed versus aesthetic depth, and autonomous physics engine scripting.
Lore & references
- The “Price Reversal” Phenomenon: Illustrates a known dynamic where a model with a significantly lower token list price ($4/$20 for Opus 5.5 vs. $10/$50 for Astra) ends up being drastically more expensive overall ($14.28 vs. $2.20) due to deeper reflection, multi-step debugging, and verbose output token generation.
- Claude Code vs. Codex CLI: Highlights the ongoing ecosystem rivalry between Anthropic’s terminal agent and OpenAI’s terminal CLI harness.
Visual style & craft
- Layout: A clean, structured side-by-side benchmark dashboard with synchronized 3D viewport recordings above a data comparison table.
- Visual generation: The simulations are rendered in 3D WebGL/canvas scenes coded by the models. Opus 5.5 prioritized richer environment modeling (textured wooden walls, lighting, polished surfaces), while GPT-6 Astra wrote a streamlined, minimalist, functional wireframe structure that solved the objective with fewer tokens and in a fraction of the time.
Described by gemini-3.8-flash on 2026-10-10 from the video’s audio and frames.
Made by AI
- Made with
- Claude Opus 5.5, GPT-6 Astra
- How we know
- X post (2026-09-22): ‘Both built a working Rube Goldberg machine with physics from scratch’
- Human role
- Same prompt given to both models; creator posted a split-screen comparison.
- Pipeline
- Opus 5.5 and GPT-6 Astra each code a physics simulation → split-screen recording
- Series
- Code-rendered film (LLM writes the program that draws every frame)