[BREAKING] Gemini 4 Argon Arrives! #1 in 13 out of 19 Official Benchmarks, Google Strikes Back!
AI時短ラボ · 2026-09-30 · review · 26,981 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
This Japanese explanatory video, presented by the synthesized Voicevox characters Zundamon and Shikoku Metan on the channel AI時短ラボ, breaks down Google DeepMind's official announcement of Gemini 4 Argon (published September 30, 2026). The presenters review the model's benchmark performance across knowledge work, coding, STEM, and cybersecurity, its 1-million-token output capability, pricing, internal Google deployments, and safety evaluations.
What is shown
- [00:00] Title screen and overview introducing Gemini 4 Argon and its competition against GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5.5.
- [00:17] Chapter 1: Details of the official Google DeepMind blog post by Koray Kavukcuoglu, rollout schedule, and the Fairwind cyber defense program.
- [01:23] Chapter 2: Pricing graphics showing input/output rates, prompt cache discounts, and the new 1-million-token output context window.
- [02:08] Chapter 3: Official comparison table for knowledge work benchmarks (Vals Index, AutomationBench, Vals Finance Agent v2, Harvey's Legal Agent Benchmark).
- [03:13] Chapter 4: Bar chart of DeepSWE v1.1 coding benchmark, followed by a comprehensive table showing FrontierSWE v2, Vibe Code Bench, Terminal-Bench 4.0, and PostTrainBench.
- [04:34] Chapter 5: Benchmark tables covering science/math (Terminal-Bench Science Q3, LABBench 2, RiemannBench), long-context GraphWalks, computer use (Agent's Last Exam, OSWorld 2.0), multimodal comprehension (Chartography, LVBench), and cybersecurity.
- [05:42] Chapter 6: Discussion of four internal deployment use cases inside Google (quantum computing optimization, datacenter memory reclamation, C/C++ to Rust migration, and video decoder acceleration).
- [06:50] Chapter 7: Cybersecurity bar charts comparing Argon against Gemini 3.8 Flash Cyber on Real-world Vulnerability Discovery and Wiz Penetration Test Benchmark, as well as the CWE-bench v1 leaderboard.
- [08:22] Chapter 8: Gray Swan IPI benchmark chart assessing susceptibility to prompt injection attacks across frontier models, and explanations of Google's four safety measures.
- [09:50] Chapter 9: Full summary comparison table compiling all 19 official benchmark metrics.
Claims & numbers
- The presenters say the announcement was published on Google's official blog at 5:00 AM JST on September 30, 2026, authored by Google DeepMind's Koray Kavukcuoglu [00:19].
- Initial access is restricted to vetted cyber defenders via the "Fairwind" program under U.S. government pre-deployment testing frameworks; wider release will follow for paid API users and Google AI Ultra subscribers [00:45, 01:10].
- The presenter states introductory pricing is $2.00 per 1M input tokens and $10.00 per 1M output tokens (with a 95% discount for cached inputs), rising to $4.00 input and $20.00 output after the introductory period [01:26].
- Maximum output limit is expanded from 64,000 tokens to 1,000,000 tokens [01:48].
- Knowledge work scores stated:
- Vals Index: Gemini 4 Argon 68.9%, GPT-6 Astra 63.1%, Claude Opus 5.5 67.0%, Claude Fable 5.1 65.8% [02:26].
- Zapier AutomationBench: Argon 51.3%, Opus 5.5 42.5%, Astra 41.4%, Fable 5.1 31.4% [02:30].
- Vals Finance Agent v2: Argon 65.4%, Astra 53.5%, Fable 5.1 58.9%, Opus 5.5 58.6% [02:46].
- Harvey's Legal Agent Benchmark: Argon 19.6%, Astra 5.4%, Fable 5.1 6.7%, Opus 5.5 3.8% [02:54].
- Coding scores stated:
- DeepSWE v1.1: Argon 77.9%, Opus 5.5 74.2%, Astra 74.1%, Fable 5.1 67.4% [03:20].
- FrontierSWE v2: Astra 65.5%, Opus 5.5 62.3%, Fable 5.1 56.3%, Argon 55.0% (Argon ranked 4th) [03:47].
- Terminal-Bench 4.0: Opus 5.5 66.4%, Astra 58.2%, Fable 5.1 57.9%, Argon 57.4% [03:57].
- PostTrainBench: Opus 5.5 49.3%, Argon 45.3%, Astra 44.3%, Fable 5.1 40.2% [04:08].
- Vibe Code Bench: Argon 91.9%, Opus 5.5 90.3%, Fable 5.1 90.3%, Astra 89.6% [04:23].
- Science, context, and multimodal scores stated:
- Terminal-Bench Science Q3: Astra 68.1%, Argon 57.6% [04:40].
- LABBench 2: Argon 88.8%, Astra 85.4% [04:51].
- RiemannBench: Argon 76.0%, Astra 72.0% [04:54].
- GraphWalks (256k to 1M tokens): Argon 84.2%, Astra 71.8%, Opus 5.5 66.8%, Fable 5.1 65.0% [05:01].
- Agent's Last Exam: Argon 39.5%, Astra 34.2% [05:16].
- OSWorld 2.0: Astra 72.6%, Argon 69.2% [05:22].
- LVBench (long video): Argon 91.7%, Astra 87.5% [05:29].
- Chartography: Argon 71.6%, Astra 71.0% [05:34].
- Internal Google use claims:
- Quantum computing: solved component optimization in minutes, surpassing published benchmarks by 40% [05:52].
- Datacenter memory: automated agent freed over 300 TiB, projecting 500 TiB to 1 PiB total memory savings across datacenters [06:01].
- Code migration: translated C/C++ repositories exceeding 800,000 lines into Rust [06:21].
- Video decoding: rewrote a 32,000-line decoder, delivering 2.7x faster performance than prior Rust code with identical output [06:37].
- Cybersecurity claims:
- Real-world Vulnerability Discovery (across 20 languages): Argon scored 85.8% vs. 71.0% for Gemini 3.8 Flash Cyber [07:11].
- Wiz penetration testing benchmark: Argon achieved 70.9% vs. 58.2% for Gemini 3.8 Flash Cyber [07:22].
- CWE-bench v1: Argon tied for #1 at 68% with Grok 4.7 and GPT-6 Astra, followed by Opus 5.5 at 67% [07:34].
- Wiz used Argon to identify a critical medical records data-leak vulnerability in widespread hospital software [08:10].
- Safety and prompt injection claims:
- Gray Swan IPI (15-attack success rate): Argon had the lowest attack success rate at 0.7%, followed by Opus 5.5 (1.0%), Fable 5.1 (1.0%), Astra (8.5%), and GPT-6 Sol (10.1%) [08:55].
- Overall benchmark record stated: Argon ranked #1 in 13 out of 19 reported categories alone, and tied for #1 in 1 category (14 total), while Astra won 3 and Opus 5.5 won 2 [09:54].
Notable quotes
- [00:00] "ついにGemini 4が来たのだ!!名前はArgon。" (Gemini 4 is finally here!! Its name is Argon.)
- [01:46] "もう1つの大きな変化が、出力の上限なのだ。前の6万4千トークンから、100万トークンに広げたのだ。" (Another major change is the output ceiling. It expanded from the previous 64,000 tokens to 1,000,000 tokens.)
- [09:52] "Gemini 4 Argonは、比較表の19項目のうち13項目で単独の1位、1項目で同率の1位なのだ。" (Out of 19 items on the comparison table, Gemini 4 Argon took sole first place in 13 items and tied for first place in 1 item.)
Assessment
This is a third-party news analysis and benchmark review summarizing Google DeepMind's official launch blog post using Voicevox character avatars. The video does not show hands-on live prompts, instead faithfully reviewing and critiquing the official graphs, benchmark tables, and claims released by Google.
Described by gemini-3.8-flash on 2026-10-01 from the video's audio and frames.