Gemini 4 Argon Is Google's Most Powerful AI Model + Early Tests!
WorldofAI · 2026-09-30 · review · 94,897 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
In this video, the creator behind the YouTube channel WorldofAI reviews Google DeepMind's newly announced frontier model, Gemini 4 Argon. The presenter breaks down its technical specifications, benchmark results, pricing structure, and architecture details, while showcasing several early community test demos including complex 3D WebGL/Three.js interactive web apps and SVG vector graphics.
What is shown
- 00:05 – 00:39: Google’s announcement post detailing Gemini 4 Argon's initial deployment to cyber defenders via the Fairwind Program, ahead of wider availability to API users and Google AI Ultra subscribers.
- 00:40 – 01:15: Leaderboard and pricing spec cards highlighting introductory pricing ($2.00/1M input, $10.00/1M output, 95% caching discount) and a 1M-token output generation limit (expanded from 64k).
- 01:37 – 03:07: Official benchmark graphs:
- DeepSWE v1.1: Gemini 4 Argon scoring 77.9%, leading GPT-6 Astra (74.1%), Claude Fable 5.1 (67.4%), and Claude Opus 5.5 (74.2%).
- Vals Index: 68.9% overall score across professional tasks.
- Vals Finance Agent v2: 65.4% score.
- Harvey's Legal Agent Benchmark: 19.6% score.
- AutomationBench: 51.3% score.
- 03:16 – 04:45: Artificial Analysis charts showing cost-efficiency ($1.99 per Intelligence Index task), an Intelligence Index score of 53 (matching GPT-6 Astra), and a recorded 15% hallucination rate.
- 06:09 – 06:45: "Sakura Pagoda Sanctuary" voxel demo generated with 106,788 voxels, featuring real-time lighting changes (day, sunset, night) and camera orbit controls.
- 06:46 – 07:23: An SVG render test producing a detailed vector graphic of a PS5 DualSense controller.
- 07:24 – 08:06: Side-by-side floatplane physics simulation test comparing Gemini 4 Pro against Claude Opus 5.5.
- 08:07 – 08:35: "Hollow & Hearth" interactive Halloween Three.js character diorama showing dynamic lighting, materials, and character controls.
- 08:36 – 09:14: "Apis Mechanica" interactive 3D mechanical bee demo with exploded view inspection, flight animations, and technical callouts.
- 09:15 – 10:02: Detailed 3D architectural model of the Roman Colosseum with section cutaways, day/night cycles, and explanatory annotations.
- 10:03 – 10:35: "Apex Fab" semiconductor wafer fabrication cleanroom interactive 3D viewer featuring exploded components and technical specs.
- 10:36 – 11:17: Four-way 3D generation comparison of a knitted Halloween bouquet across Sonnet 5.5 Max, Gemini 4 Pro, GPT 6.1 Sol Max, and Opus 5.5 Max.
- 11:18 – 11:52: "Cloudline" interactive railway world demo comparing GPT-6 Astra against Gemini 4 in 3D scene construction and UI.
Claims & numbers
- Availability: The presenter notes Argon is currently restricted to cyber defenders and trusted testers in Google DeepMind's Fairwind Program, with paid API and Google AI Ultra rollouts expected soon.
- Pricing:
- The presenter states an introductory price of $2.00 per 1M input tokens and $10.00 per 1M output tokens, with cached input discounted by 95% ($0.10/1M tokens).
- Post-introductory pricing is cited as $4.00 per 1M input tokens and $20.00 per 1M output tokens.
- Token Output Window: Output token limit increased from 64k to 1,000,000 tokens in a single generation trajectory.
- Benchmark Scores (DeepMind):
- DeepSWE v1.1: 77.9%.
- Vals Index: 68.9%.
- Vals Finance Agent v2: 65.4%.
- Harvey's Legal Agent Benchmark: 19.6%.
- AutomationBench: 51.3% (and 77.5% on automation workflows).
- Artificial Analysis Metrics:
- Intelligence Index score: 53 (tied with GPT-6 Astra, 1 point behind GPT-6.1 Sol).
- Task cost: $1.99 per task (~60% of Astra's task cost).
- Hallucination rate: 15%, reported as the lowest measured on the index for high-intelligence tier models.
- Model Size: Citing a Reuters report, the presenter states Gemini 4 Argon is larger in physical parameter count than previous Google "Pro" models.
Notable quotes
- 00:43 – 00:49: "This may be the best cost-to-performance model that we have ever seen."
- 01:08 – 01:15: "Google has increased the maximum output from 64k to 1 million tokens output."
- 04:33 – 04:37: "Argon has a 15% hallucination rate, the lowest on Artificial Analysis has ever measured..."
Assessment
This is a third-party commentary and summary review video aggregating official announcements from Google DeepMind, benchmark charts from Artificial Analysis, and leaked early test demonstrations shared in online community Discords. The presenter demonstrates no live prompting himself; instead, he relies on pre-recorded screen captures and community-sourced WebGL/code generation demos.
Described by gemini-3.8-flash on 2026-10-01 from the video's audio and frames.