GPT-6.1 Sol vs Claude Sonnet 5.5 – This Was NOT Even CLOSE!
Bijan Bowen · 2026-10-01 · review · 50,702 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
In this video, presenter Bijan Bowen conducts a direct head-to-head comparison between Anthropic’s Claude Sonnet 5.5 and OpenAI’s GPT-6.1 Sol across four complex multimodal and agentic benchmarks. Operating both models at maximum reasoning effort, Bowen evaluates their capabilities in game recreation, product software and hardware co-founding, physical RC car control, and audiovisual media production.
What is shown
- Introduction & Methodology [00:12]: Overview of testing methodology comparing Claude Sonnet 5.5 (Max / Ultracode) against GPT-6.1 Sol (Max / Ultra) on identical real-world prompts.
- Test 1: Demolition Derby Game Recreation [04:19]:
- Prompt requiring models to replicate a reference image of a demolition derby using Blender and Godot Engine, incorporating soft-body deformation physics, detached car parts, and multiple camera views [05:12].
- GPT-6.1 Sol’s game test run in Godot [06:26] and its acknowledgement of taking shortcuts on rigid chassis approximations [09:32].
- Claude Sonnet 5.5’s game test run in Godot, featuring more faithful vehicle contours, customized menus, HUD, and destruction mechanics [11:38].
- Test 2: Chatbot Hardware & Branding Overhaul ("Talk Buddy") [16:24]:
- Prompt providing existing 3D CAD files, build manuals, and hardware specs for a physical desktop robot, tasking each model to build complete firmware using OpenAI's Realtime Voice API, branding, marketing website with 3D exploded render previews, a 15-second promo video, and installation scripts [18:54].
- Sonnet 5.5’s created brand "Peri", featuring a comprehensive interactive marketing website with WebGL product renders [21:48], brand guidelines [27:52], a 15-second launch commercial [29:05], and live hardware interaction testing on the physical robot [29:32].
- GPT-6.1 Sol’s created brand "Mote", featuring its website [31:51], downloadable press kit [34:31], launch video [34:51], and live device test [35:32].
- Side-by-side comparison of Mote vs. Peri [37:48].
- Test 3: Autonomous RC Car Parallel Parking [40:55]:
- Setup using a top-down camera and an Arduino Uno connected to a cloned 27 MHz RF breakout board [41:46] to control a red Tyco RC car without proportional throttle/steering [43:05].
- GPT-6.1 Sol controlling the car via camera image pulses [44:40] and generating a motion infographic video explaining its calibration and path planner [46:34].
- Claude Sonnet 5.5 executing vehicle moves [49:02], encountering a safety intervention notice ("Message Flagged" for potential knowledge distillation / reasoning extraction) when prompted for debrief graphics [51:17], and outputting its motion debrief animation [52:26].
- Test 4: Rap Promo Mixing & Video Production ("MATMUL") [53:32]:
- Audio and video assets provided: 7 stems (6 instrumental, 1 raw vocal recorded on iPhone at 85 BPM), cover art, and raw b-roll clips [55:47].
- GPT-6.1 Sol's final mixed audio and promotional video [60:57].
- Claude Sonnet 5.5's final mixed audio, autotune processing, and video edit with on-screen kinetic typography and screen tracking [63:55].
- Conclusion & Verdict [67:58]: Final breakdown of results, with the presenter expressing surprise that Sonnet 5.5 significantly outperformed GPT-6.1 Sol across creativity, detail, and execution.
Claims & numbers
- The presenter notes both models are available under comparable $200/month subscription tiers [70:52].
- For Test 1, GPT-6.1 Sol ran for just under 1 hour and 40 minutes to build the playable Godot prototype [06:20].
- During the 3D rendering phases of Test 2, the presenter notes his laptop reached temperatures between 99°C and 102°C [29:28].
- For Test 3, the radio controller operates on an Arduino Uno over 27 MHz with non-proportional, binary pulses (full-throttle/steering on or off) [43:05].
- Sonnet 5.5 was briefly blocked by a safety classifier when the presenter requested a debrief of its parking logic, citing potential knowledge/reasoning distillation [51:17].
- For Test 4, the track tempo is 85 BPM, consisting of 6 instrumental tracks and 1 vocal track recorded on an iPhone [56:22].
- GPT-6.1 Sol took 54 minutes on Ultra to complete the music mix and video render [60:44], while Claude Sonnet 5.5 took approximately 1 hour 50 minutes to 2 hours [63:23].
- The presenter claims Sonnet 5.5 consistently produced higher quality, less "apathetic" results across coding, motion graphics, and audio mastering compared to GPT-6.1 Sol [70:04].
Notable quotes
- [01:20]: "In the testing videos that I did and put on this channel, I found Sonnet 5.5 was mopping the floor with GPT-6.1 Sol, which for some reason almost made me upset because it just felt wrong."
- [26:46]: "I mean, that right there... this, I would have no issue putting this live right now."
- [70:34]: "I found that Sonnet 5.5 wiped the floor with Sol. It may have taken three times as many tokens and been more expensive, but my specific interest in this video is which of these models is better."
Assessment
This is a hands-on review and empirical benchmark comparison conducted by an independent developer evaluating frontier agentic workflows. All coding, rendering, hardware interfacing, and media edits are demonstrated on screen with source files and terminal outputs visible, though longer generation times are compressed for video pacing.
Described by gemini-3.8-flash on 2026-10-02 from the video's audio and frames.