Gemini 4 Argon
Sam Witteveen · 2026-10-01 · review · 29,478 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Sam Witteveen breaks down Google DeepMind's announcement of Gemini 4 Argon, the first model in the Gemini 4 family. Through animated diagrams and benchmark charts from Artificial Analysis, he examines its intelligence rankings, 1-million-token output window, pricing structure, hallucination rates, and agentic coding performance.
What is shown
- [00:10] Timeline graphic showing Gemini 4 Argon announced on September 30, currently in testing with selected users, and coming soon to the public.
- [00:30] Diagram contrasting Argon's 1-million-token single-response output ceiling against the 128K limits of other frontier models.
- [00:58] Illustrated explanation of model verbosity and token inefficiency seen in Gemini 3.6 Flash and 3.7 Flash.
- [01:20] Roadmap noting the 7-month gap between Gemini 3.1 Pro Preview and Gemini 4 Argon.
- [01:40] Artificial Analysis Intelligence Index chart displaying Argon scoring 53, tied with GPT-6 Astra and ahead of GPT-6.1 Sol (52).
- [03:00] Schematic showing how chunking and context compaction cause detail loss compared to Argon's 1-pass full-output capability.
- [04:30] Google case study slide on porting
libgav1to safe Rust via iterative agent profiling and rewriting, resulting in a 2.7x speedup. - [05:41] Artificial Analysis chart of output tokens per task, followed by task-cost comparisons.
- [07:51] AA-Omniscience non-hallucination chart highlighting Argon's 85% rate (15% hallucination rate) versus Astra and Sol.
- [08:33] AutomationBench-AA and Terminal-Bench 4.0 leaderboard comparisons against Claude Sonnet 5.5, Claude Opus 5.5, and GPT-6 Astra.
Claims & numbers
- Gemini 4 Argon was announced on September 30, 2026, and is initially available only in testing to selected users (the presenter says).
- It can output up to 1,000,000 (1M) tokens in a single response, roughly 8x higher than the 128K ceiling of models like Claude Opus 5.5, Claude Fable 5.1, and GPT-6 Astra (the presenter says).
- On the Artificial Analysis Intelligence Index, Argon scores 53, matching GPT-6 Astra at maximum reasoning, edging out GPT-6.1 Sol (52), and jumping 23 points above Gemini 3.1 Pro Preview (30) (the presenter says).
- It marks Google's first Pro/frontier release in 7 months since 3.1 Pro Preview (the presenter says).
- In a Google case study, Argon agents ported the
libgav1video decoder to safe Rust, achieving a 2.7x speedup with identical output (the presenter says). - At an estimated 100 tokens/second, generating 1M tokens would take approximately 3 hours of wall-clock decoding time (the presenter says).
- Gemini 3.8 Flash generation speed is noted at around 300 tokens/second (the presenter says).
- Launch pricing for Argon is $2/million input tokens and $10/million output tokens (a 50% launch discount off the standard $4 / $20), with cached inputs discounted by 95% (the presenter says).
- Argon costs an estimated $1.99 per task on the Artificial Analysis Intelligence Index—about 60% of GPT-6 Astra ($3.26), but roughly 2.7x the cost of GPT-6.1 Sol ($0.72) (the presenter says).
- Without the price discount, Argon uses approximately 1.2x the tokens per task compared to Astra (the presenter says).
- On the AA-Omniscience benchmark, Argon demonstrates an 85% non-hallucination rate (only 15% made up), compared to 51% made up for Astra and 54% made up for GPT-6.1 Sol, though its lower coverage puts overall accuracy at 50% (overall Omniscience score 42 vs. Astra's 43) (the presenter says).
- Argon ranks #1 on AutomationBench-AA at 77.5%, outperforming Claude Sonnet 5.5 (71%) and Claude Opus 5.5 (70%) (the presenter says).
- On Terminal-Bench 4.0, Argon scores 57%, trailing Claude Sonnet 5.5 (64%), Claude Opus 5.5 (60%), and GPT-6 Astra (60%) (the presenter says).
Notable quotes
- [00:01] "Okay, so Google has finally announced Gemini 4. So, this is actually called Gemini 4 Argon."
- [01:57] "This is 23 points ahead of Gemini 3.1 Pro Preview, which was their last frontier Pro model."
- [06:23] "Fewer tokens per tasks means lower cost and lower latency."
Assessment
This is an independent analysis and commentary video by an AI educator/reviewer discussing Google's announcement and third-party benchmark evaluations from Artificial Analysis. No live interactive prompting or raw model access is shown; all discussions are based on Google's blog disclosures and Artificial Analysis leaderboard data.
Described by gemini-3.8-flash on 2026-10-01 from the video's audio and frames.