GPT-6 SOL vs Luna vs Claude Opus 5.5: Which Should You Use?
AI with Surya · 2026-09-23 · review · 17,345 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
In this hands-on benchmark review, Surya (from the channel AI with Surya) compares Anthropic’s Claude Opus 5.5 against OpenAI’s GPT-6 Sol and GPT-6 Luna following their simultaneous launch on September 22, 2026. Using a custom local benchmarking tool called "Model Arena" connected via OpenRouter, he runs all three models side-by-side across three front-end coding challenges of increasing complexity to assess generation speed, token cost, thinking behavior, and code quality.
What is shown
- [00:00 - 02:23] Context overview presenting launch-day announcements, API pricing charts ($0.50 to $20/1M output tokens), Artificial Analysis Intelligence Index scores, and AutomationBench task completion figures.
- [02:24] Introduction of the custom "Model Arena" dashboard running on
localhost:3000, measuring time, thinking tokens, output tokens, and dollar cost for each model side-by-side. - [02:46] Test 1 Prompt: Generating a single-file HTML landing page for an umbrella brand named "Squall," requiring animated wind/rain resistance, feature breakdowns, testimonials, and pre-order pricing.
- [04:13 - 06:12] Test 1 evaluation: GPT-6 Luna finishes first in 1m 23s ($0.0063), GPT-6 Sol in 1m 48s ($0.013), and Claude Opus 5.5 in 3m 25s ($0.52). Surya tests each rendered page full-screen, highlighting Opus 5.5's dynamic canvas storm and gust animations.
- [06:47] Test 2 Prompt: Generating an interactive fleet operations dashboard tracking 12 delivery trucks navigating coastal storm bands, featuring live route disruption and rerouting buttons.
- [07:20 - 10:35] Test 2 evaluation: Luna finishes in 1m 41s ($0.0075) and Sol in 1m 49s ($0.012). Sol successfully calculates vehicle avoidance routes while Luna's trucks remain stranded. Claude Opus 5.5 finishes in 9m 17s ($1.27) after 30k thinking tokens, rendering an operations console with multi-layer radar heatmaps and status tracking.
- [10:44] Test 3 Prompt: Building a self-contained 3D browser sailing game called "Storm Run" with Three.js/WebGL, navigational buoys, stormy ocean waves, lightning, and rogue wave hazards.
- [11:09 - 15:10] Test 3 evaluation: Luna generates a basic, barely functional 3D canvas (rated 3/10) in 1m 35s ($0.0075); Sol produces a playable 3D sailboat game with checkpoints and hazard warnings in 2m 17s ($0.15); Opus 5.5 finishes in 18m 07s ($2.41, using 74.5k thinking tokens and 123.1k output tokens), generating a photorealistic storm game complete with dynamic wave crests, physics, lighting, and an interactive rogue wave sequence.
Claims & numbers
- Pricing & generation differences: The presenter states that GPT-6 Luna costs roughly 40x less per output token than Claude Opus 5.5 ($0.50 vs $20 per 1M output tokens) [00:46]. Claude Opus 5.5 is priced at $4 input / $20 output per 1M tokens (reported ~40% cheaper than Opus 5) [01:09]. OpenAI cut GPT-6 Sol ($2 / $10) and Luna ($0.10 / $0.50) prices roughly in half compared to GPT-5.6 Sol and Luna [01:21].
- Benchmarks cited: Artificial Analysis Intelligence Index scores shown place Claude Opus 5.5 at 58, GPT-6 Sol at 48, and GPT-5.6 Sol at 47 [01:30]. AutomationBench business completion rates place Opus 5.5 at 40%, Sol at 33%, and Luna at ~21% [02:04].
- Live Arena test metrics:
- Test 1 (Landing Page): Luna (1m 23s, 745 thinking tokens, 12.5k output tokens, $0.0063); Sol (1m 48s, 495 thinking tokens, 12.4k output tokens, $0.013); Opus 5.5 (3m 25s, 856 thinking tokens, 26.1k output tokens, $0.52) [04:14, 05:58].
- Test 2 (Operations Dashboard): Luna (1m 41s, 2.2k thinking tokens, 14.5k output tokens, $0.0075); Sol (1m 49s, 2.2k thinking tokens, 12.4k output tokens, $0.012); Opus 5.5 (9m 17s, 30.0k thinking tokens, 63.7k output tokens, $1.27) [07:24, 09:31].
- Test 3 (3D Game): Luna (1m 35s, 3.2k thinking tokens, 14.8k output tokens, $0.0075); Sol (2m 17s, 2.3k thinking tokens, 14.6k output tokens, $0.15); Opus 5.5 (18m 07s, 74.5k thinking tokens, 123.1k output tokens, $2.41) [11:15, 11:23].
Notable quotes
- [00:46] "The cheapest of the three, Luna, costs about 40x less than Opus 5.5."
- [01:46] "Nobody seems to be pacing the price cuts."
- [15:31] "As long as you don't have a very complicated task, I think you can easily go with Luna and save a ton of money and still get the job done."
Assessment
An authentic, independent benchmark and review demonstrating live model outputs through OpenRouter API calls. While extended generation wait times are edited down for pacing, the live code outputs, token metrics, and interactive browser executions are genuine, thoroughly tested, and honestly critiqued.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.