Claude Sonnet 5.5 a TERMINÉ OpenAI : Claude est devenu cheaté
Melvynx · 2026-09-29 · community · 9,948 views · Français
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
In this video, French software developer and AI educator Melvynx reviews Anthropic’s newly released Claude Sonnet 5.5 alongside Claude Opus 5.5. He analyzes Artificial Analysis benchmark figures and runs side-by-side evaluations across interactive 3D physics, technical educational apps, and motion graphics video generation against OpenAI's GPT-6 Astra and GPT-6 Sol.
What is shown
- Artificial Analysis Benchmarks [01:03]: Melvynx walks through Excalidraw slides displaying the Artificial Analysis Intelligence Index and Coding Agent Index, highlighting Claude Code with Sonnet 5.5 scoring 68 and Opus 5.5 scoring 66, ahead of GPT-6 Astra.
- Cost, Speed, and Token Efficiency Comparisons [02:29]: Charts detailing cost per intelligence index task, speed/latency per task, and output token usage across thinking budget settings (Medium vs. xHigh/Max).
- Local Benchmark Dashboard [06:01]: Melvynx showcases a custom testing interface (
localhost:9080) tracking automated model execution across complex coding tasks completed on September 28, 2026. - Interactive 3D Car Crash Simulation [08:52]: Side-by-side evaluation of Three.js/physics implementations. Opus 5.5 Medium creates an interactive 3D simulation with wall destruction and vehicle replay controls [09:07], whereas Opus 5.5 xHigh stalls [09:47], Sonnet 5.5 Medium/xHigh has glitchy collision physics [10:04], GPT-6 Astra crashes/loads flat [11:21], and GPT-6 Sol fails completely [11:45].
- 3D Air Conditioning Explanatory App [12:40]: Opus 5.5 Medium generates a detailed, animated interactive 3D house model demonstrating refrigerant loops and heating/cooling mechanics [12:45], compared against Sonnet 5.5 [14:02] and GPT-6 Astra's static 2D illustration [14:47].
- Motion Graphics Video Generation Benchmark (Lumail Ad) [23:15]: Playback of 45-second HTML/canvas motion graphic marketing videos for email tool "Lumail". Opus 5.5 xHigh produces a polished, timed product video with typography and interface animations [23:15], Sonnet 5.5 xHigh produces a functional but visually disjointed rendition [24:03], and GPT-6 Astra generates a flat, non-animated dark mockup [26:19].
- Workflow Recommendations [27:00]: Melvynx outlines practical guidelines for choosing thinking effort budgets, recommending Opus 5.5 at Medium for standard tasks and reserving xHigh only for complex architectural tasks.
Claims & numbers
- Benchmark Scores: The presenter states Claude Code with Sonnet 5.5 achieves a top score of 68 on the Artificial Analysis Coding Agent Index, outperforming Opus 5.5 (66) and GPT-6 Astra (62, 6 points lower) [01:45].
- Thinking Budget Costs: The presenter claims running Sonnet 5.5 at Max thinking budget costs up to $7.60 per task compared to $3.46 for Opus 5.5 xHigh on benchmarked tasks [02:35], but Sonnet 5.5 on Medium drops to around $0.60 per task while retaining solid capability [03:33].
- Execution Times: The presenter claims Sonnet 5.5 Medium is significantly faster than GPT-6 Astra Medium and Opus 5.5 Medium on standard tasks [03:45].
- Run Cost Discrepancy: In his custom benchmark runs, Melvynx notes that Opus 5.5 Medium cost $7.71 over ~49 minutes [09:22], whereas Opus 5.5 xHigh cost $16.39 over 1 hour 38 minutes [08:41] while delivering worse physics results.
- Switching Cost Philosophy: The presenter claims subscription switching costs between AI vendors are negligible ("costs nothing"), arguing developers should opportunistically change tools based on who currently holds the performance crown [28:28].
Notable quotes
- "Sonnet 5.5 vient de sortir et il est meilleur que Opus 5.5, qui est lui-même meilleur que Astra..." [00:00]
- "En réalité, en fait, quand je regarde ici, on peut voir que le Medium a mieux fonctionné que le Extra High, hein." [09:03]
- "Opus 5.5 est actuellement le OG... Utilisez Opus 5.5 Medium pour la majorité des tâches." [27:00]
Assessment
This video is an independent review and hands-on benchmark evaluation by an AI developer. The demonstrated applications and web apps are shown live inside browser tabs, showcasing both the successes of Claude Opus 5.5/Sonnet 5.5 at medium reasoning effort and the diminishing returns or regressions observed when pushing thinking budgets to maximum levels.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.