Sol 6.1 shouldn't be so cheap
Theo - t3.gg · 2026-09-29 · review · 32,666 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Theo Browne (t3.gg) reviews OpenAI’s newly released GPT-6.1 Sol, comparing its benchmark results, cost efficiency, and practical coding performance against GPT-6 Astra, GPT-6 Sol, and Anthropic’s Claude Opus 5.5 and Sonnet 5.5. He analyzes its steep price cuts and aggressive caching discounts, tests it on real-world repository audits and codebase reviews, and demos "fishslop," a 3D web game generated by the model.
What is shown
- Benchmark Comparisons: Terminal-Bench 4.0 score vs. cost per task plot ([00:38], [03:22], [08:47]) and DeepSWE score vs. cost per task plot comparing GPT-6.1 Sol, GPT-6 Astra, Claude Opus 5.5, and Jev Router ([04:18], [19:34]).
- Sponsor Segment: Depot and Depot Metal for accelerated CI and container builds ([01:37]–[03:21]).
- Excalidraw Notes: Diagramming GPT-6.1 Sol's traits, token pricing, and caching math ([05:21], [07:10], [10:00], [12:40]).
- X (Twitter) Posts: Tibo’s post announcing changes to the $200 Pro subscription ([05:54]) and Theo’s previous post illustrating response quality consistency between GPT-6 Astra and Claude Fable 5.1 ([10:09]).
- Audit Benchmarks & Tables:
- Orchestrator V2 audit cost, score, and token telemetry comparing GPT-6.1 Sol, Astra, and Opus 5.5 ([13:22]).
- T3 Code fleet audit score table showing GPT-6.1 Sol scoring 87.4 ([14:12]).
- Claude Opus 5.5 reviewing GPT-6.1 Sol's performance and guessing its pricing ([15:40]–[18:59]).
- 3D Game Demo ("fishslop"): A 3D underwater submarine and fish game built with WebGL/Three.js by GPT-6.1 Sol, demonstrating high-quality 3D plant and submarine models alongside cluttered, low-quality UI and sluggish movement ([21:18]–[26:06]).
- Real-World Code Review: Auditing GitHub PR #11836 in T3 Code to catch state desync and worktree setup regressions ([26:07]–[28:05]).
Claims & numbers
- The presenter says Anthropic released Claude Opus 5.5 the previous week, prompting OpenAI to quickly launch GPT-6 Sol and GPT-6 Luna, which he found unimpressive ([00:08]).
- GPT-6.1 Sol API pricing is stated as $2.00 per million input tokens, $10.00 per million output tokens, and $0.10 per million cached input tokens (a 95% cache discount, down from the standard 90%) ([07:10], [07:30]).
- The presenter notes Tibo's post indicates that OpenAI's revamped $200 Pro subscription yields roughly half the dollar API spend equivalent compared to the previous plan ([06:12]).
- On DeepSWE, the presenter reports GPT-6.1 Sol on low scored 80% (16/20 tasks) at $0.21 per task in 4.8 minutes, matching Claude Opus 5.5 on max (80% at $14.65 per task in 46.8 minutes) and Astra on low (80% at $1.46 per task in 4.6 minutes) ([04:37], [21:09]).
- On Terminal-Bench 4.0, the presenter claims GPT-6.1 Sol scored 65.1%–66.7% on max runs at $1.38 per task ([03:54], [09:32]).
- In a multi-model audit of T3 Code, GPT-6.1 Sol placed first with an 87.4 score costing ~$2.15, ahead of Astra (83.8), Grok 4.7 (80.7), and Opus 5.5 (79.9 at $5.00) ([14:12]–[14:38]).
- The presenter claims GPT-6.1 Sol costs about 1/5th the price of GPT-6 Astra and roughly 70x cheaper than Opus 5.5 for equivalent DeepSWE benchmarks ([05:01], [07:21]).
- The presenter states generating the "fishslop" 3D project cost roughly $5 ([25:42]).
Notable quotes
- "This model on low costs 21 cents, versus Astra on low costing a dollar forty-six, and Opus 5.5 on max getting the same score for 14 dollars and 65 cents." [04:37]
- "This model's token price is cheaper than Sonnet, but its token utilization is still maintaining OpenAI's usual efficiency, which results in just crazy price to performance." [14:40]
- "This model is unacceptably garbage at UI. It has regressed again." [24:43]
Assessment
This is an independent hands-on technical review by Theo Browne evaluating pre-release and launch access to OpenAI's GPT-6.1 Sol. The video shows real benchmark dashboards, live code audits in private repos, and an interactive browser-based 3D demo while openly discussing both the model's strengths in auditing/pricing and its shortcomings in frontend UI and long-running autonomous development.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.