Gemini 4 Argon, for the little there is to say
Salvatore Sanfilippo · 2026-10-01 · review · 31,428 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Salvatore Sanfilippo (antirez) reviews Google DeepMind's announcement blog post for Gemini 4 Argon, analyzing its reported benchmark results, token limits, and pricing. He discusses the competitive dynamics between major US frontier AI labs and Chinese open-weights developers, arguing against vendor lock-in as model leadership continues to shift.
What is shown
- [00:00] Screen recording of the Google DeepMind blog post announcement: "Gemini 4 Argon: our next era of frontier intelligence" (dated September 30, 2026).
- [06:30] The blog post text detailing the initial restricted rollout through the Fairwind Program, pricing ($2 per million input tokens and $10 per million output tokens, 95% discount on cached inputs), and internal engineering applications at Google.
- [07:27] The comprehensive benchmark comparison table contrasting Gemini 4 Argon with GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5.5 across coding, math, agentic reasoning, long context, and cybersecurity benchmarks.
- [09:33] The section highlighting Gemini 4 Argon's generation capabilities, specifically expanding output tokens up to 1 million tokens (up from 64k).
- [11:18] CWE-bench v1 cybersecurity defense leaderboard and mitigation charts.
- [11:47] Sections describing frontier safeguards, sandboxing, and rollout schedule via the Fairwind Program and Google AI Ultra subscribers.
- [13:06] Sanfilippo navigates to Google One to check the Google AI Ultra subscription tier pricing (€99.99/mo base, €219.99/mo for the 30 TB tier).
- [13:46] Demonstration of his custom-built local screen and webcam recording tool.
Claims & numbers
- The presenter notes that OpenAI recently restructured its high-end plans, cutting usage allowances on its $200 tier and introducing a $500 monthly tier.
- The presenter highlights Google's introductory pricing for Gemini 4 Argon: $2 per million input tokens, $10 per million output tokens, with a 95% discount on cached tokens (effective $0.10/M tokens).
- The presenter points out Gemini 4 Argon's 1 million output token generation limit, contrasting continuous generation capacity with input context windows.
- The blog post and presenter state benchmark figures comparing Gemini 4 Argon to rivals:
- Vais Index: 68.9% (vs. GPT-6 Astra 63.1%, Claude Fable 5.1 65.8%, Claude Opus 5.5 67.0%).
- AutomationBench: 51.3% (vs. Astra 41.4%, Fable 5.1 31.4%, Opus 5.5 42.5%).
- DeepSWE v1.1: 77.9% (vs. Astra 74.1%, Fable 5.1 67.4%, Opus 5.5 74.2%).
- Vibe Code Bench: 91.9% (vs. Astra 89.6%, Fable 5.1 90.3%, Opus 5.5 90.3%).
- CWE-bench v1: 68.0% (vs. Astra 66.0%, Fable 5.1 58.0%, Opus 5.5 67.0%).
- The presenter notes that Claude Opus 5.5 outperforms Argon on Terminal Bench 4.0 (66.4% vs 57.4%) and PostTrainBench (49.3% vs 45.3%), while GPT-6 Astra leads on Terminal-Bench Science 0.1 (68.1% vs 57.6%) and OSWorld-2.0 (72.6% vs 69.2%).
Notable quotes
- [01:28] "Non c'è la salsa magica, la ricetta speciale... questi sono gli ingredienti."
- [05:30] "Non mi stancherò mai di dire che chi si lega a uno specifico provider... si perde in questo momento... la possibilità di pagare i token il meno possibile col modello migliore possibile."
- [09:59] "Se io domani arrivo dal nulla e ho il modello che è migliore a tutti gli altri, ho vinto. Gli altri se ne possono andare a casa."
Assessment
This is an independent commentary and reaction video by a well-known software engineer reading and evaluating Google's official blog post announcement. The presenter does not run live model evaluations himself, instead reviewing Google's published numbers and pricing while offering industry commentary on provider lock-in and model capabilities.
Described by gemini-3.8-flash on 2026-10-02 from the video's audio and frames.