I Put GPT-6 Sol and Opus 5.5 to the Test: Here's What Happened
Eric Tech · 2026-09-23 · review · 12,531 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
In this video, creator Eric (Eric Tech) conducts a side-by-side benchmark comparison between OpenAI's GPT-6 Sol and Anthropic's Claude Opus 5.5 across multiple development and agent tasks. He tests both models on fixing a minor CSS bug, implementing a complex chart feature in a production financial web app, building a 3D Chongqing open-world browser game, running an autonomous web-search and computer-use rental lead research task, and generating an interactive 3D travel globe application.
What is shown
- [00:00] Intro displaying OpenAI's GPT-6 Sol / Luna launch page alongside Anthropic's Claude Opus 5.5 announcement page (dated September 22, 2026).
- [00:32] Test 1 (Small Bug): Both models fix a dialog alignment bug in Eric's production app Finfluencer. Both succeed; GPT-6 Sol finishes faster (4 min, 60k tokens) than Opus 5.5 (6 min, 75k tokens).
- [02:53] Test 2 (Big Bug / Feature Addition): Implementing interactive avatar selection linking creators to stock timeline points on a Tesla chart. Sol 6 generates a working, clean UI implementation in 9m 32s using ~40k tokens, beating Opus 5.5 (17m 3s, 234k tokens).
- [06:44] Test 3 (Chongqing 3D Game): Testing browser-based 3D playable games built by both models. Opus 5.5's build (Mountain City Chongqing, port 5190) includes custom audio, police AI with wanted levels, pedestrian interactions, and minimap, while Sol 6's version (The City Has Layers, port 5188) lacks audio, combat interaction, and has broken collision geometry.
- [10:45] Test 4 (Computer Use / Rent Scan): Running deep research in EricOS for Vancouver apartment rentals. Sol 6 uses browser/computer vision tools to inspect images and listings, completing in 9m 11s and returning 8 deduplicated, verified listings. Opus 5.5 deploys 37 sub-agents, consuming ~4.12M tokens over 45 minutes, returning ~200 mostly unverified/duplicate listings without visual validation.
- [14:58] Test 5 (3D Travel Globe): Comparing Sol 6's app (Atlas, port 4173) and Opus 5.5's app (Wayfarer, port 5173). Opus 5.5's build features animated flight paths, camera transitions, and procedural 3D city buildings (Dubai, Tokyo) with weather data, judged superior in UX and visual quality despite taking longer (35m 26s vs 13m 35s).
- [19:23] Final summary scorecard reviewing all five categories: Sol 6 wins in token efficiency, speed, small bug fixing, and computer use; Opus 5.5 wins in game development and 3D visual application design.
Claims & numbers
- Small Bug Fix: The presenter reports Claude Opus 5.5 consumed 75k tokens and took 6 minutes, while GPT-6 Sol consumed 60k tokens and took 4 minutes.
- Feature Addition (Big Bug): The presenter shows terminal logs indicating Opus 5.5 used 234,429 tokens and 114 tool calls across 17 minutes 3 seconds, whereas GPT-6 Sol took 9 minutes 32 seconds and ~40,000 tokens.
- 3D Game Generation: The presenter shows Opus 5.5 took 1 hour 15 minutes 33 seconds and 472k tokens, whereas GPT-6 Sol took ~50 minutes and 745,939 tokens.
- Autonomous Rental Research (Computer Use): The presenter shows GPT-6 Sol took 9 minutes 11 seconds to find 8 verified listings; Claude Opus 5.5 took ~45 minutes and 4,004,923 tokens across 37 sub-agents and 342 tool calls, yielding ~245 raw records that were mostly duplicates.
- 3D Globe Application: The presenter reports Opus 5.5 (Wayfarer) took 35 minutes 26 seconds and ~200k tokens, while GPT-6 Sol (Atlas) took 13 minutes 35 seconds and 141,137 tokens.
Notable quotes
- [02:34] "Definitely I would say that Sol, GPT-6 here definitely wins on this one."
- [10:20] "Overall though, I definitely think that results matter, because especially for building a game here, user experience here definitely count first."
- [20:00] "In terms of specifically fixing bugs, get to the straight points, I definitely feel like GPT-6 Sol here is definitely better for that."
Assessment
This is an authentic third-party technical review and live screen demonstration comparing local Vite dev builds generated by GPT-6 Sol and Claude Opus 5.5. The tests, terminal execution logs, token counts, and interactive browser applications are shown running directly on the host machine without deceptive staging.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.