GPT-6 Sol VS Opus 5.5 (Fully Tested): I DID A SIDE-BY-SIDE Comparison of BOTH MODELS!
AICodeKing · 2026-09-23 · review · 20,549 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
In this review video, AICodeKing presents a side-by-side benchmark comparison between OpenAI’s GPT-6 Sol and Anthropic’s Claude Opus 5.5, both released on September 22, 2026. The presenter analyzes vendor specs and public benchmarks before running both models through his proprietary 8-task "KingBench 3" evaluation and four larger "Long Horizon" app-building tests using his "Bambood" coding harness.
What is shown
- [00:08] Side-by-side display of the launch announcements for GPT-6 Sol and Claude Opus 5.5.
- [02:08] Comparison slides detailing standard API token pricing, cache read pricing, and context window limits for both models.
- [02:41] Artificial Analysis Intelligence Index v4.3.2 scores and cost-per-task metrics compared on bar charts.
- [03:29] Public benchmark scores compared across Terminal-Bench 4.0, SciCode, Humanity’s Last Exam, and AutomationBench-AA.
- [04:16] Demonstration of the presenter's testing environment ("Bambood"), running local coding sessions with Codex and Claude Code CLI tools.
- [04:52] KingBench 3 Task 1: Interactive elevator simulation test; Opus scores 8/10, Sol scores 7/10.
- [05:30] KingBench 3 Task 2: Interactive 3D contact lens case; Opus scores 10/10 with detailed lenses inside, while Sol scores 6/10 due to cap clipping issues.
- [06:05] KingBench 3 Task 3: Interactive 3D folding table with slider control; Opus scores 8/10, Sol scores 7/10.
- [06:26] KingBench 3 Task 4: SVG generation of a panda eating a burger; both receive 10/10.
- [06:36] KingBench 3 Task 5: 2D bow and arrow target archery game; Opus scores 9/10, Sol scores 6/10 due to basic mechanics and lack of curved trajectories.
- [07:07] KingBench 3 Task 6: Combinatorics calculation (target answer: 20,460); both models score 10/10.
- [07:13] KingBench 3 Task 7: Panda fine-tuning workflow (creating dataset, fine-tuning Gemma 2B, and building a local UI); both score 10/10.
- [07:45] KingBench 3 Task 8: Interactive 3D wristwatch with live dual timezone displays; both score 10/10.
- [08:08] Final KingBench 3 scoreboard and updated leaderboard showing Opus 5.5 taking #1.
- [08:33] Long Horizon KingBench demonstrations of four complex apps:
- [08:46] Terminal Movie Tracker using TMDB API (Sol unfinished; Opus fully functional).
- [09:20] A4 Poster Studio integrating Fal API and 3D preview (Opus visually preferred).
- [09:53] 3D interactive Blu-ray shelf application (Opus produced richer physics and spine details).
- [10:34] Markdown note-taking workspace with integrated OpenCode agent (Opus produced a more complete UI).
Claims & numbers
- The presenter states that both GPT-6 Sol and Claude Opus 5.5 were released on September 22, 2026.
- The presenter states GPT-6 Sol API pricing is $2.00 per million input tokens and $10.00 per million output tokens (50% cheaper than GPT-5.6 Sol promotional rates), with context caching reads at $0.20 per million tokens and an input surcharge above 272K tokens.
- The presenter states Claude Opus 5.5 API pricing is $4.00 per million input tokens and $20.00 per million output tokens, with cached input reads at $0.20 per million tokens.
- The presenter notes both models feature ~1M context token windows (Sol specified at 1.05M) and a 128K maximum output token limit.
- On Artificial Analysis Intelligence Index v4.3.2, the presenter reports:
- Medium effort: Sol scores 40, Opus 5.5 scores 51.
- Max effort: Sol scores 48, Opus 5.5 scores 58.
- Cost per task: Sol costs $0.25 (medium effort) vs. $1.34 for Opus 5.5 (~5.4x cost difference).
- On individual benchmarks reported by Artificial Analysis at medium effort:
- Terminal-Bench 4.0: Opus 5.5 scores 53% vs. Sol 19%.
- SciCode: Opus 5.5 scores 59% vs. Sol 54%.
- Humanity’s Last Exam: Opus 5.5 scores 55% vs. Sol 41%.
- AutomationBench-AA: Opus 5.5 scores 61% vs. Sol 58%.
- In the presenter's KingBench 3 (8 tasks at medium effort):
- GPT-6 Sol scored 66/80 (82.5%).
- Claude Opus 5.5 scored 75/80 (93.75%).
- On the presenter's KingBench 3 leaderboard: Opus 5.5 ranks #1 (93.75%), followed by Fable 5.1 (92.5%), GLM 5.3 (91.25%), GPT-6 Astra (90%), and GPT-6 Sol tied with Fable 5 at 82.5%.
- The presenter claims Opus 5.5 won all four of his qualitative Long Horizon app builds.
Notable quotes
- [02:05] "For the API, Opus costs $4 per million input tokens and $20 per million output tokens. So Sol's standard input and output rates are half the price."
- [08:00] "Sol gets 66 out of 80, which is 82.5%. Opus gets 75 out of 80, which is 93.75%. That's a lead of 11.25 percentage points for Opus."
- [11:10] "I kept getting results that felt more complete, with more attention paid to the details I would otherwise have to fix myself."
Assessment
This is an independent user review and hands-on developer benchmark comparing real outputs from two AI models inside coding and app development environments. The demonstrations show real code execution and interactive web applications, though scoring on KingBench 3 and the long-horizon builds reflects the creator's subjective evaluation of code and UI completeness.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.