GPT-6 Sol i Opus 5.5: Szum vs Rzeczywistość [Test agentów i recenzja]
SmartTech Synergy · 2026-09-27 · review · 17,712 views · Polski
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
In this review video, a presenter from the Polish tech channel SmartTech Synergy evaluates and compares two recently released frontier AI models: OpenAI’s GPT-6 Sol and Anthropic’s Claude Opus 5.5. He analyzes their technical specifications, pricing, and independent benchmark scores before running a hands-on coding agent comparison where both models build a full-stack image processing web application from scratch.
What is shown
- [00:22] - [01:01]: Overview slides contrasting the model hierarchies of OpenAI (GPT-6 Astra, Sol, Luna) and Anthropic (Opus 5.5, Fable 5.1, Opus 5), alongside the Artificial Analysis Intelligence Index leaderboard.
- [01:02] - [03:11]: GPT-6 Sol technical and pricing overview slides showing API rates, context window, knowledge cutoff, and benchmark tables (including GDPval and HealthBench).
- [04:03] - [06:33]: Claude Opus 5.5 overview slides showing pricing, context limits, effort settings, and benchmark scores across Terminal-Bench 4.0, FrontierCode 1.1, AutomationBench, and AA-Briefcase.
- [07:07] - [07:49]: The test task specification: mockups and requirements for "ClearCut AI", a web application requiring background removal (using local neural models like BiRefNet and RMBG-1.4 with GPU acceleration and CPU fallback), cropping/rotation tools, responsive UI, and Tinyfy compression.
- [07:53] - [10:27]: Claude Opus 5.5 execution in Claude Code: deep initial research into ONNX runtimes and GPU vs. CPU execution benchmarks, agentic implementation, and real-time automated browser verification.
- [10:35] - [11:46]: Demonstration of the completed app generated by Opus 5.5, showcasing pixel-accurate frontend fidelity, functional background removal, interactive cropping, and mobile responsiveness.
- [12:06] - [14:19]: GPT-6 Sol execution in OpenAI Codex: workspace leakage incident, automated testing issues with external Chrome, and visual inspection of the resulting app (which lacked GPU inference, broke the source selector, and had a lagging crop tool).
- [15:37] - [17:12]: GPT-6 Sol attempting bug fixes, hitting the 5-hour quota limit (93% consumed), and leaving the application incomplete.
Claims & numbers
- GPT-6 Sol:
- The presenter notes API pricing is $2.00 / 1M input tokens and $10.00 / 1M output tokens ($0.20 cache read, doubling above a 272k token prompt threshold), representing a 50% price cut compared to GPT-5.6 Sol.
- Context window is 1,050,000 tokens with a maximum output of 128,000 tokens; knowledge cutoff is April 20, 2026.
- On the Artificial Analysis Intelligence Index, GPT-6 Sol scores 48 points (versus 47 for GPT-5.6 Sol), while task execution cost dropped ~47% from nearly $2.00 to $1.06 per task.
- In GDPval-AA v2.1, Sol dropped approximately 100 points compared to its predecessor (scoring 1487 vs. 1588 for GPT-5.6 Sol).
- HealthBench Professional score is 60.8 (compared to 60.5 for GPT-5.6 Sol).
- Claude Opus 5.5:
- The presenter reports API pricing is $4.00 / 1M input tokens and $20.00 / 1M output tokens ($0.20 cache read; no long-context surcharge), making it 20% cheaper than Opus 5 and 60% cheaper than Fable 5.1.
- Context window is 1,000,000 tokens with 128,000 max output; knowledge cutoff is June 2026.
- Takes 1st place on the Artificial Analysis Intelligence Index with 58 points (compared to 51 for Opus 5 and 53 for Fable 5.1).
- Benchmark scores shown: GDPval-AA (1844), AA-Briefcase v1.1 (1822), Terminal-Bench 4.0 (66.4), FrontierCode 1.1 (54.6), AutomationBench (42.5), Agents' Last Exam (63.2).
- Agent Test Results:
- Claude Opus 5.5 completed the full production-grade application in 49 minutes, consuming 39% of a 5-hour Pro subscription limit and 6% of the weekly limit.
- GPT-6 Sol spent 28.5 minutes on its first pass and an additional 28.5 minutes attempting repairs (57 minutes total), exhausting 93% of the 5-hour limit and 14% of the weekly limit while delivering an incomplete and partially broken application.
Notable quotes
- [00:10]: "Co do tego ostatniego okazało się bzdurą, wiemy już, że nie są, ale obydwie premiery są ciekawe. Choć jedna bardziej." ("As for the latter, it turned out to be nonsense; we already know they aren't, but both releases are interesting. Though one more so.")
- [10:09]: "A teraz nie mam żadnych wątpliwości, że Opus zrobi to lepiej i szybciej." ("And now I have no doubt that Opus will do it better and faster.")
- [14:43]: "Krótko mówiąc, to że jest tańszy od Opusa 5.5 w API, zupełnie nie przekłada się na to, ile możemy z nim zrobić w agencie." ("In short, the fact that it is cheaper than Opus 5.5 in the API does not translate at all into how much we can do with it in an agent.")
Assessment
This is an authentic, independent third-party hands-on benchmark and review video. The presenter clearly shows the setup, prompt specifications, live terminal logs, web browser test interactions, and resulting codebases, offering a fair and transparent comparison of both models running in realistic agent environments without visible misleading cuts.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.