As of: 2026-10-10 23:43 CEST. Researched and written by AI agents (Claude Opus 5.5 in Claude Code). Human editor: Adam Bicz. Canonical page: https://postcutoff.com/v/leogao25-opus-5-5-vs-gpt-6-astra-rube-goldberg-cost-x/ # Opus 5.5 vs GPT-6 Astra build the same physics Rube Goldberg machine: cost and time compared @leogao25, 22 September 2026, X. 1,420 views as of 10 October 2026. Kind: Made by AI. Watch: https://x.com/leogao25/status/2102544078927741369 ## Why it is here Curated in github.com/athemeroy/awesome-claude-5-5-videos ('Mixed workflows and comparisons'). Post claims Opus 5.5 cost $14.28 and 62 min vs GPT-6 Astra $2.20 for the same prompt, despite lower per-token price (single user's test, unverified). ## Description (written by Gemini from the video) **Summary** This side-by-side comparison video, posted by user @leogao25 on September 22, 2026, compares Anthropic's Claude Opus 5.5 and OpenAI's GPT-6 Astra coding and executing a self-running 3D physics Rube Goldberg machine from scratch. Both models were given a single high-effort run from an empty directory using their respective developer harnesses (Claude Code vs. Codex CLI), with the screen showing the resulting 3D simulations alongside performance, token, and cost metrics. --- **What is shown** * **Prompt & benchmark setup**: Text at the top states: *"Opus 5.5 vs GPT-6 Astra: same prompt, same machine"* with the prompt: *"Build a self-running Rube Goldberg machine. Physics from scratch. One attempt per model, high effort."* [00:00]. * **Left pane (Claude Opus 5.5 via Claude Code)**: * [00:00–00:06]: A metal sphere drops down a multi-level wooden ramp within a detailed wooden room/cabinet enclosure. * [00:07–00:10]: The ball rolls across a wooden tabletop, striking and toppling a line of dominoes. * [00:11–00:16]: The falling domino triggers a counterweight pulley system that raises a checkered finish flag up a flagpole. * **Right pane (GPT-6 Astra via Codex CLI)**: * [00:00–00:06]: A marble rolls through elevated green curved wireframe tracks on a minimalist green plane. * [00:07–00:10]: The marble trips a row of standing upright pegs and a seesaw rocker lever. * [00:11–00:16]: The kinetic impulse triggers a pulley contraption that hoists an orange triangular target flag. * **Comparative performance table**: Detailed breakdown showing build time, API cost, output token count, and pricing tiers [00:00–00:16]. --- **Claims & numbers** * **Claude Opus 5.5**: * Harness: Claude Code. * Build time: 62 minutes. * API cost: $14.28. * Output tokens: 265K. * List price per million tokens: $4 input / $20 output. * **GPT-6 Astra**: * Harness: Codex CLI. * Build time: 11 minutes. * API cost: $2.20. * Output tokens: 20K. * List price per million tokens: $10 input / $50 output. * **Ratio (Opus 5.5 relative to Astra)**: * Build time: 5.8× longer. * API cost: 6.5× higher. * Output tokens: 13.6× more. * List price ratio: 0.4× (cheaper list price per token). * **Run conditions**: Stated at the bottom: *"Each model started in an empty folder with no config files. Cost is API list price for every token used, including the agent's pre-training. Time is wall-clock from prompt to finished file. One run each."* --- **Notable quotes** * *"Build a self-running Rube Goldberg machine. Physics from scratch. One attempt per model, high effort."* [00:00] * *"Each model started in an empty folder with no config files."* [00:00] * *"Time is wall-clock from prompt to finished file. One run each."* [00:00] --- **Assessment** This is a real head-to-head developer benchmark of two frontier models building complex interactive simulations autonomously. While both successfully produced functioning physics contraptions from a cold start, the test reflects a single trial (`One run each`) and compares two distinct agent harnesses (Claude Code vs. Codex CLI), which significantly influences token volume, reasoning loops, and wall-clock time alongside raw model capability. --- **Lyrics & themes** * The video is completely silent with no audio, narration, or soundtrack. * Key themes focus on real-world coding agent cost efficiency, speed versus aesthetic depth, and autonomous physics engine scripting. --- **Lore & references** * **The "Price Reversal" Phenomenon**: Illustrates a known dynamic where a model with a significantly lower token list price ($4/$20 for Opus 5.5 vs. $10/$50 for Astra) ends up being drastically more expensive overall ($14.28 vs. $2.20) due to deeper reflection, multi-step debugging, and verbose output token generation. * **Claude Code vs. Codex CLI**: Highlights the ongoing ecosystem rivalry between Anthropic's terminal agent and OpenAI's terminal CLI harness. --- **Visual style & craft** * **Layout**: A clean, structured side-by-side benchmark dashboard with synchronized 3D viewport recordings above a data comparison table. * **Visual generation**: The simulations are rendered in 3D WebGL/canvas scenes coded by the models. Opus 5.5 prioritized richer environment modeling (textured wooden walls, lighting, polished surfaces), while GPT-6 Astra wrote a streamlined, minimalist, functional wireframe structure that solved the objective with fewer tokens and in a fraction of the time. _Described by gemini-3.8-flash on 2026-10-10 from the video's audio and frames._ ## Made by AI - Made with: Claude Opus 5.5, GPT-6 Astra - How we know: X post (2026-09-22): 'Both built a working Rube Goldberg machine with physics from scratch' - Human role: Same prompt given to both models; creator posted a split-screen comparison. - Pipeline: Opus 5.5 and GPT-6 Astra each code a physics simulation → split-screen recording - Series: Code-rendered film (LLM writes the program that draws every frame) ## Related - 2026-09-22: [Anthropic releases Claude Opus 5.5](https://postcutoff.com/e/2026-09-22-claude-opus-5-5/)