GPT-6.1 Sol Is HERE – Can THIS Beat Claude Opus 5.5?
Bijan Bowen · 2026-09-29 · review · 21,229 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
In this video, creator Bijan Bowen benchmarks OpenAI's newly released GPT-6.1 Sol against a suite of intensive coding, physical robotics, and 3D simulation tasks. Across multiple tests—including complex browser-based games, Godot/Blender projects, physical robot arm control, and legacy hardware troubleshooting—Bowen assesses whether GPT-6.1 Sol delivers on its promise of near-Astra capabilities at one-fifth the cost, comparing its performance to Anthropic's Claude Sonnet 5.5 and Claude Opus 5.5.
What is shown
- [00:05] OpenAI Announcement Page & Specs: Overview of GPT-6.1 Sol pricing ($2.00/1M input, $10.00/1M output), 1,000,000-token context window, 128k maximum output, and DeepSWE benchmark graphs comparing Sol to GPT-6 Sol and GPT-6 Astra.
- [03:01] ChatGPT Pro Plan Usage: Checking the account usage dashboard, showing 98% of the weekly limit remaining on the $200/month plan before running tests.
- [03:21] Backyard Pool Party (Godot/Blender Game): Testing an isometric diving mini-game generated in Godot and Blender, played in real-time, followed by a rerun at max reasoning effort [09:47] that finishes at [10:30] and is tested at [12:09].
- [04:44] Single-Script Browser OS ("Halo OS"): Opening a self-contained HTML OS containing functional 3D mini-games ("Signal City" driving game at [06:01] and "Orbital Run" at [07:22]), an email client [07:50], notes app [08:14], synth sound lab [08:28], and "Continuum" state capsule saver [09:15].
- [13:20] Standalone C++ NYC Skateboarding Game: Generating an OpenGL/C++ street skateboarding game ("Borough - NYC 2002") featuring a custom cinematic replay mechanism when pressing 'Z' [15:11], followed by an upgraded revision [16:53] with improved textures, water rendering, and board grinding [18:00].
- [19:02] Physical Robot Arm Manipulation: Testing an embodied robotic arm tasked with manipulating a toy car; the model calculates coordinates for 18.5 minutes before stopping and reporting it cannot safely secure a grip [20:31].
- [20:39] 3D Subway FPS Scene ("Last Line"): Running a Three.js survival horror shooter set in an underground subway station, showing weapon mechanics, zombie waves [21:26], dynamic lighting, and riding the subway train between stations [22:14].
- [24:04] 3D Seinfeld Apartment Replica: Generating an interactive Three.js walkthrough of Jerry Seinfeld's apartment [25:54], featuring accurate room layouts, easter egg props [26:33], lighting controls [27:50], and dollhouse views [27:27].
- [28:21] Old School RuneScape PvP Replication: Inspecting a browser reproduction of RuneScape's Grand Exchange PvP area [28:49], showing authentic UI menus, inventory items, combat mechanics, and sound effects [29:38].
- [32:04] Interactive 3D Saab 900 Turbo Model: Loading a detailed WebGL/Three.js car visualizer [32:15] allowing camera rotation, color swapping, opening the hood to view the modeled engine bay [32:23], interior inspection [34:06], and opening doors/hatch [34:25].
- [34:53] Legacy Laptop Setup & Daybreak Mode: Attempting to install Linux onto an old Pentium III Gateway laptop using an Ethernet PCMCIA card and OpenAI's Daybreak model [35:52] with visual monitoring.
- [37:17] Final Resource Usage & Conclusion: Verifying that running all tests consumed only 3% of the weekly plan limit (dropping from 98% to 95%).
Claims & numbers
- Pricing & Positioning: The presenter highlights that GPT-6.1 Sol is priced at $2.00 per million input tokens and $10.00 per million output tokens, which OpenAI positions as "near-Astra intelligence for a fifth of the price" ($10/$50 on Astra) and directly matches Claude Sonnet 5.5's pricing.
- Model Parameters & Limits: Displays a 1,000,000 context window, a 128,000 max output token limit, and an April 30, 2026 knowledge cutoff.
- Benchmarks: DeepSWE benchmark graphs shown on screen indicate GPT-6.1 Sol achieves 75.2% on high reasoning effort ($0.65 cost per task) and 71.9% on max reasoning effort ($1.57 cost per task).
- Token Efficiency: The entire suite of heavy coding and multimodal tasks across the video only decreased the presenter's weekly ChatGPT Pro quota by 3% (from 98% to 95%).
Notable quotes
- [01:05] "Really this model is competing in price with Claude Sonnet 5.5, which as we saw yesterday is a surprisingly, surprisingly capable model."
- [18:00] "Check this out. That is absolutely like X-Games mode, pro skate... that was awesome."
- [38:15] "I don't want to make a very definitive stance on this model based off of a lot of visual 3D tasks, but I'll say I'm not as impressed as I was hoping to be, especially following the test of Sonnet 5.5."
Assessment
This is an authentic, hands-on independent review and technical stress test by YouTuber Bijan Bowen. The demonstrations—spanning web apps, native C++ executables, Godot rendering, and physical hardware control—are run live on screen without pre-rendered trickery, complete with the presenter candidly showcasing both model successes and failures.
Described by gemini-3.8-flash on 2026-09-30 from the video's audio and frames.