I Gave Claude Opus 5.5 a full set of house plans. Did it follow them?
The AI Essentials · 2026-09-28 · review · 5,698 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Justin Geis from The AI Essentials reviews and tests Anthropic's Claude Opus 5.5 model, focusing on its performance in 3D modeling tasks. He evaluates its benchmark improvements and pricing before demonstrating its capabilities via MCP (Model Context Protocol) integration in Blender and SketchUp, comparing results against OpenAI's GPT-6 Astra.
What is shown
- [00:16] Anthropic's announcement page for Claude Opus 5.5, detailing performance benchmarks, pricing, and coding agent capabilities.
- [03:08] A 3D modeling test prompt using a multi-pass instruction structure (overall form, detail refinement, and final inspection) with reference images for an Eames lounge chair and ottoman via a Blender MCP server.
- [03:32] Side-by-side visual comparison in Blender between models created by GPT-6 Astra and Claude Opus 5.5, evaluating geometry, mesh smoothness, materials, and adherence to reference images.
- [07:20] The "Farmhouse test" feeding complete architectural plan drawings from FreeFarmhouse.com to Opus 5.5 to generate an accurate 3D model in SketchUp.
- [08:04] Dimension verification showing interior layout accuracy and dimension drift in the GPT-6 Astra model versus Claude Opus 5.5.
- [11:59] Claude Opus 5.5 generating a self-audited dimension discrepancy table comparing drawing dimensions against model dimensions.
- [12:47] SketchUp/LayOut output where Claude Opus 5.5 automatically generated drawing overlay checks against the 3D model, as well as an exported multi-page architectural presentation plan set with site plans, exterior elevations, and floor plans.
Claims & numbers
- The presenter notes Claude Opus 5.5 was released on September 22, 2026.
- Quoting Anthropic's published pricing table, Opus 5.5 costs $0.20 per million cache read tokens, $4 per million input tokens, $20 per million output tokens, and $5 per million cache write tokens (compared to Opus 5 at $0.50, $10, $50, and $12.50 respectively).
- The presenter shows Anthropic's benchmark table where Opus 5.5 scores 66.4% on Terminal-Bench 4.0 (versus Fable 5.1 at 55.8%, GPT-6 Astra at 57.9%, and GPT-5.6 Sol at 37.3%) and 67.7% on Humanity's Last Exam (compared to 64.9% for Fable 5.1 and 67.2% for GPT-6 Astra).
- The presenter claims Opus 5.5 adhered significantly closer to exact blueprint dimensions than Astra, often within 1/16th of an inch of specified dimensions, though it ran slower than Astra.
Notable quotes
- [04:52] "While it did a better job of creating the model itself, it didn't do as good of a job following the reference image..."
- [11:22] "So I mean overall, I would say that this is doing a better job of paying attention in the long run."
- [13:14] "And so that was super cool. But then the other thing it did, which I did not expect and I didn't even know that it could do, is it also created a bunch of LayOut views..."
Assessment
This is an independent hands-on review and practical workflow evaluation by a 3D modeling creator. The tests are executed in real software (Blender, SketchUp, and LayOut) using MCP integrations, showing both the strengths (blueprint adherence, automated LayOut sheet creation) and visible imperfections (rough meshes, misplaced doors, and small dimension errors).
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.