The Sol 6.1 Benchmarks Are STUPID, So I Tested It vs Sonnet 5.5
Chase AI · 2026-09-29 · review · 128,362 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Chase AI reviews OpenAI's newly announced GPT-6.1 Sol following OpenAI DevDay, analyzing its benchmark scores and head-to-head performance against Anthropic's Claude Sonnet 5.5 across four real-world coding and generation challenges. Through side-by-side evaluations of automated canvas explainer videos, frontend web design, interactive 3D web graphics, and a browser-based 3D tank game, the presenter evaluates whether GPT-6.1 Sol's dramatic token efficiency and lower task costs offset Claude Sonnet 5.5's advantages in speed and visual design execution.
What is shown
- [00:00] OpenAI DevDay announcement and benchmark slides: OpenAI's launch page for GPT-6.1 Sol, showing DeepSWE cost vs. score graphs and OSWorld 2.0 offline set comparisons between GPT-6.1 Sol, GPT-6 Sol, and GPT-6 Astra.
- [01:13] Benchmark breakdown: Artificial Analysis index comparisons across Intelligence Index, Finance & Accounting, Strategy & Ops, Legal, Healthcare & Medical, Engineering, and Economics, alongside lab benchmarks (DeepSWE v1.1 and AutomationBench 1.0.6) and effort-level task costs.
- [03:23] Test 1 — JavaScript Explainer Video: Prompting both models to generate a self-contained 15-second vertical animated canvas video explaining Claude subagents with synchronized audio. GPT-6.1 Sol's output is reviewed ([03:48]), followed by Claude Sonnet 5.5's output ([04:20]).
- [04:56] Test 2 — Landing Page Design ("Dune House"): Prompting both models to create a responsive landing page for a fictional boutique hotel named "Dune House". GPT-6.1 Sol generated images natively ([06:01]), while Claude Sonnet 5.5 generated images via a Higgsfield MCP integration ([05:08], [07:17]).
- [08:28] Test 3 — 3D Interactive Travel Globe: Testing Three.js interactive data visualizations. GPT-6.1 Sol created "ORBIT" ([08:28]), an interactive 3D globe with flight routes, while Claude Sonnet 5.5 produced "MERIDIAN" ([09:33]), featuring audio, cinematic fly-in camera animations, and destination cards.
- [10:55] Test 4 — 3D Browser Game: Prompting both models to clone a World of Tanks-style game using Three.js. GPT-6.1 Sol produced "IronFront: Armored Warfare" ([11:04]), and Claude Sonnet 5.5 generated "Iron Vanguard" ([12:06]), both showing playable tanks, aiming, HUDs, and destructible elements.
Claims & numbers
- The presenter says GPT-6.1 Sol was announced at OpenAI DevDay immediately following the underwhelming release of GPT-6.
- On DeepSWE v1.1, the presenter states GPT-6.1 Sol scored 75.2% (High effort, $0.65 cost per task), beating GPT-6 Astra's 73.2% (Max effort, $7.50 cost per task) and Sonnet 5.5's score by +4.2 points.
- On AutomationBench 1.0.6, the presenter reports Sonnet 5.5 leads Sol by +8.7 points (44.8% max with fallbacks vs. 36.1% max).
- On token API list pricing, both models share identical rates: $2 per million input tokens and $10 per million output tokens.
- The presenter claims Sonnet 5.5 generates tokens roughly twice as fast as GPT-6.1 Sol (138 tokens/second vs. 69 tokens/second).
- The presenter shows Artificial Analysis cost-per-task data indicating Sol is up to 10x cheaper at maximum effort ($0.72 vs. $7.60), 7x cheaper at extra high ($0.39 vs. $2.74), 3.4x cheaper at high ($0.32 vs. $1.08), and 2.8x cheaper at medium ($0.21 vs. $0.59); when normalized to 52% benchmark accuracy, Sol was 4x cheaper.
- The presenter states OpenAI recently slashed usage allowances on its $200/month plan by 50% (from 20x to 10x), while noting Anthropic's 20x tier had a 5-hour rolling limit rather than an uncapped weekly allowance.
- Explainer video test token usage: Sol used 70,000 tokens versus Sonnet 5.5's ~300,000 tokens.
- Landing page test token usage: Sol used 109,000 tokens versus Sonnet 5.5's ~200,000 tokens.
- 3D globe test token usage: Sol used ~150,000 tokens versus Sonnet 5.5's ~400,000 tokens.
- 3D tank game test: Sol took over 2 hours and ~350,000 tokens; Sonnet 5.5 took ~1 hour and ~700,000 tokens.
Notable quotes
- [02:39]: "Basically, Sol: super, super, super token-efficient."
- [07:05]: "Which probably is a good thing, because this model is so dang cheap that if this is giving me Astra performance, what's there to complain about?"
- [13:28]: "6.1 Sol kind of feels like just a slight step below Sonnet 5.5, yet it's way more efficient with tokens. It's cheaper, but it's slower."
Assessment
This is a hands-on review and comparative benchmark demonstration by an independent AI creator. The presenter directly demonstrates genuine browser-rendered outputs (HTML5 canvas, Three.js applications, and front-end layouts) and provides full transparency regarding token consumption, runtime, and qualitative shortcomings for both models.
Described by gemini-3.8-flash on 2026-10-01 from the video's audio and frames.