Gemini 4 Argon explained in 5min..
Caleb Writes Code · 2026-10-01 · review · 135,473 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
In this video, Caleb from the YouTube channel "Caleb Writes Code" breaks down Google's release of its flagship model, Gemini 4 Argon. He analyzes its benchmark performance—particularly on agentic coding—evaluates benchmark contamination issues, and compares Google's tooling and business strategy against Anthropic and OpenAI.
What is shown
- [00:00] Timeline diagram depicting the gap between Gemini 3.1 Pro and Gemini 4 Argon, noting the release of intermediary Flash models (3.5 to 3.8 Flash).
- [00:13] Model tier lists and an X post from Elon Musk reacting to model rankings.
- [00:24] Benchmark table comparing Gemini 4 Argon against GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5.5 across knowledge work, agentic coding, ML engineering, math, and multimodal benchmarks.
- [00:34] DeepSWE score versus cost scatter plot, followed by Epoch AI's benchmark review page highlighting flaws in DeepSWE v1.1.
- [01:01] Breakdown of evaluation benchmarks into public, semi-private, and private datasets, along with an Epoch AI chart on rapid benchmark saturation.
- [01:39] Diagram contrasting models paired with their agentic coding harnesses (Gemini 4 Argon with Antigravity, Claude Opus 5.5 with Claude Code, GPT-6 with Codex).
- [01:57] Metrics showing weekly active users for OpenAI's Codex versus Google's Antigravity.
- [02:16] Artificial Analysis chart plotting Intelligence Index versus Cost per Task.
- [03:11] Slide from Google's Q2 earnings call citing token processing volume and developer adoption.
- [03:53] Visual illustration of context generation comparing a 64k token limit to Argon's 1M output tokens, explaining autoregressive error accumulation and referencing Yann LeCun's critiques.
Claims & numbers
- The presenter notes it has been about seven months between Google's prior flagship announcements and Gemini 4 Argon [00:00].
- Gemini 4 Argon scores 77.9% on DeepSWE v1.1, placing it above GPT-6 Astra (74.1%), Claude Opus 5.5 (74.2%), and Claude Fable 5.1 (67.4%) [00:30].
- Epoch AI found that 23 out of 113 tasks (over 20.3%) in the DeepSWE benchmark were flawed [00:41].
- Codex has over 5 million weekly active users, compared to 2.4 million weekly active users for Google's Antigravity platform [01:58].
- Argon launches at an introductory API price of $2 per million input tokens and $10 per million output tokens, increasing to $4 input and $20 output after the promotional period [02:11].
- According to Alphabet's Q2 earnings call, over 9 million developers build with Google models monthly, and the API processes approximately 22 billion tokens per minute (up from 16 billion tokens per minute the prior quarter), surpassing 11 quadrillion tokens annualized on API channels alone [03:11].
- Gemini 4 Argon expands the maximum output token limit from 64,000 to 1 million output tokens [03:56].
- In autoregressive generation with 99% per-step reliability across 100 independent steps, theoretical compound accuracy drops to roughly 37% [04:30].
Notable quotes
- "It's been about seven months since Google announced a new flagship model Gemini 4 Argon." [00:00]
- "From a recent report, Codex has more than doubled the amount of weekly active users at 5 million compared to 2.4 million for Antigravity." [01:57]
- "Google extended their output tokens from 64,000 to 1 million tokens... definitely a huge move from Google because maintaining coherence as autoregressive models generate longer and longer tokens gets increasingly difficult..." [03:55]
Assessment
This is an independent analysis and review video examining Gemini 4 Argon's launch, benchmark integrity, and market context. The video uses published benchmark tables, news reports, financial statements, and diagrams rather than live, interactive software demonstrations.
Described by gemini-3.8-flash on 2026-10-02 from the video's audio and frames.