GEMINI 4 is nuts...
Wes Roth · 2026-10-01 · review · 52,710 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Wes Roth discusses Google’s announcement of Gemini 4 Argon and its internal engineering applications, benchmarks, and Sergey Brin’s direct involvement at Google. He also provides an overview of OpenAI’s DevDay 2026 announcements—including “dots” autonomous agents, GPT-6.1 Sol, the Ultrafast speed tier, the Decisions API, and the new $500/month Pro plan—alongside a sponsored workflow demonstration of CodeRabbit’s ChangeStack review tool.
What is shown
- [00:00] Google DeepMind’s official blog post announcing Gemini 4 Argon by Koray Kavukcuoglu, detailing pricing, rollout via the Fairwind Program, and internal Google deployment metrics.
- [03:34] Demo of CodeRabbit’s ChangeStack tool reviewing an AI-assisted pull request for a persistent video research board, displaying stack breakdown layers, semantic diffs, and an interactive logic flow diagram.
- [08:29] Detailed benchmark comparison table displaying Gemini 4 Argon against GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5.5 across reasoning, coding, science, and agentic evals.
- [10:36] Multi-benchmark evaluation charts from Artificial Analysis showing Gemini 4 Argon in high reasoning mode.
- [11:45] Business Insider report on Sergey Brin working out of a Mountain View microkitchen to guide Google’s AI strategy.
- [14:50] OpenAI DevDay 2026 launch pages, including dots personal agents, dot.com redirecting to Grok Bot [16:08], GPT-6.1 Sol API specifications [17:25], and Ultrafast mode [21:10].
- [22:36] Demo of TypeSafe AI’s Jev decision model playing Tetris, Minesweeper, and Halite via millisecond-level categorical decisions, contrasted with OpenAI's Decisions API.
- [25:20] OpenAI product pages for Agents API with computer use, AWS Bedrock integration, ChatGPT Space, MCP plugins, and ChatGPT Pro 500 tier pricing.
Claims & numbers
- Gemini 4 Argon:
- The presenter says Gemini 4 Argon launches at an introductory pricing of $2 per million input tokens and $10 per million output tokens, with cached input tokens discounted by 95% ($0.10/M).
- The presenter notes the output token limit is expanded to 1 million tokens (up from 64k).
- The presenter cites Google’s internal results: Argon optimized quantum computing spacetime resource subroutines, beating the published baseline by 40% in minutes; memory efficiency agents freed over 300 TiB across data centers with an estimated 500 TiB to 1 PiB total savings; and codebase agents rewrote 32,000 lines of SIMD C/C++ code in libgav1 into safe Rust that vectorized automatically and runs 2.7x faster than the initial Rust port.
- Benchmark scores shown for Gemini 4 Argon: Vals Index (68.9%), AutomationBench (51.3%), Vals Finance Agent v2 (65.4%), Harvey’s Legal Agent Benchmark (19.6%), DeepSWE v1.1 (77.9%), FrontierSWE v2 (55.0%), Vibe Code Bench (91.9%), Terminal-Bench 4.0 (57.4%), PostTrainBench (45.3%), Terminal-Bench Science 0.1 (57.6%), LABBench 2 (88.8%), RiemannBench (76.0%), GraphWalks (99.7%), Agent's Last Exam (39.5%), OSWorld-2.0 (69.2%), Chartography (71.6%), LVBench (91.7%), and CWE-Bench v1 (89.1%).
- Google Operations:
- The presenter mentions Business Insider cited eight current or former employees confirming Sergey Brin is actively working with leadership on Gemini model development.
- The presenter notes Google's Project Suncatcher involves testing orbital TPU data centers using a two-satellite constellation communicating via space lasers.
- OpenAI DevDay Announcements:
- GPT-6.1 Sol is priced at $2 input and $10 output per million tokens, with a 1,050,000-token context window and 128,000 max output tokens.
- Ultrafast API service tier claims speeds up to 8x faster than standard mode.
- ChatGPT Pro 500 launches at $500/month and includes Astra Ultrafast.
- Decisions API uses GPT-6 Luna to evaluate finite user-defined questions in real time.
- TypeSafe AI Jev Demo:
- Displayed average round-trip decision speed between 104 ms and 210 ms, averaging around 140 ms per placement in Tetris, and making nearly 5,000 decisions across 67 batches in Minesweeper for an estimated API cost of a few cents.
Notable quotes
- [00:00] "Well, it looks like Google is back. Rumors of its demise were greatly exaggerated."
- [14:41] "So, I'm willing to say it, Google is back, probably."
- [22:12] "Instead of answering in a language, it just answers like, 'I think it's 99% chance the answer is yes.'"
Assessment
This is an independent news commentary and analysis video featuring a third-party workflow tool sponsorship. The presenter walks through verified blog announcements, documentation, and benchmark sheets released by Google and OpenAI, alongside real on-screen software demos of CodeRabbit ChangeStack and TypeSafe AI's Jev.
Described by gemini-3.8-flash on 2026-10-01 from the video's audio and frames.