Googles New Gemini 4 Argon is Now The Worlds Smartest AI
TheAIGRID · 2026-10-01 · review · 68,506 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Andrew Black from TheAIGRID reviews Google’s newly announced Gemini 4 Argon model and analyzes third-party and official benchmark results. He assesses whether the model represents genuine frontier progress or selective "benchmaxxing," highlighting its leadership on general evaluation boards, high token output limit, and low hallucination rate, balanced against weaker web-development coding scores and employee skepticism reported in the press.
What is shown
- [00:07] Google's official benchmark comparison sheet comparing Gemini 4 Argon against GPT-6 Astra, Claude Opus 5.5, and Claude Sonnet 5.5 across agentic coding, ML engineering, mathematics, computer use, and long-context evaluations.
- [01:51] The Arena.ai overall leaderboard displaying
gemini-4-argon-highranked #1 with an Elo score of 1525 (price listed as $2 / $10). - [02:34] Vals Index leaderboard showing Gemini 4 Argon ranked #1 at 68.90% accuracy, $15.68 cost/test, and 41m 30s latency.
- [03:20] Artificial Analysis Intelligence Index chart placing Gemini 4 Argon at a score of 55, alongside an API cost/intelligence scatter plot [03:48] and pricing breakdown [04:13].
- [04:23] Arena.ai Text Arena leaderboard showing Gemini 4 Argon (High) ranked #1 with a score of 1530.
- [05:16] Code Arena WebDev leaderboard showing Gemini 4 Argon ranked #8 (score 1,673).
- [06:04] Andon Labs' Blueprint-Bench 2 chart showing Gemini 4 Argon leading with a ~0.56 score in generating floorplans from apartment interior photos.
- [06:59] A social media post quoting Google DeepMind's announcement of a 1M token output limit for Argon.
- [07:57] Artificial Analysis Hallucination Index chart showing Gemini 4 Argon (High) recording a 15% hallucination rate.
- [09:09] A Bloomberg article titled "Google Grapples With Employee Skepticism About New Gemini Model" discussing internal concerns about real-world performance versus benchmark optimization.
Claims & numbers
- The presenter says Google released Gemini 4 Argon and claims it surpasses competitor frontier models across multiple benchmark categories.
- The presenter states that on Arena.ai's overall category, Gemini 4 Argon High ranks #1 with a score of 1525 across 4,942 votes at an API cost of $2 / $10 per million tokens.
- On the Vals Index (dated Sep 30, 2026), the presenter shows Gemini 4 Argon scoring 68.90% accuracy, outpacing Claude Sonnet 5.5 (67.94%) and Claude Fable 5.1 (65.61%).
- On the Artificial Analysis Intelligence Index, Gemini 4 Argon scores 55, positioned next to Claude Opus 5.5 (56) and GPT-6 Astra (56).
- On the Artificial Analysis pricing chart, Gemini 4 Argon is listed at $1.99 per blended token metric, compared to $2.01 for GPT-6 Astra, $5.41 for Claude Sonnet 5.5, and $5.98 for Claude Opus 5.5.
- On Arena.ai's Text Arena, Gemini 4 Argon ranks #1 with an Elo score of 1530, roughly 25 points above Claude Fable 5 and Claude Opus 5.5.
- On Code Arena WebDev, Gemini 4 Argon ranks #8 with a 1,673 score, behind Claude Opus 5.5, Claude Fable 5.1, GPT-6.1 Sol, and GPT-6 Astra.
- On Blueprint-Bench 2, Gemini 4 Argon scores ~0.56, nearing the human baseline of 0.59.
- The presenter highlights Google DeepMind's announcement that Argon features a 1M token output limit, allowing it to generate up to one million tokens in a single response rather than just having a 1M context input window.
- On the Artificial Analysis Hallucination Index, Gemini 4 Argon scores a 15% hallucination rate, substantially lower than competitors shown (Claude Sonnet 5.5 at 28%, GPT-6 Astra at 51%).
- The presenter notes a Bloomberg report stating that while Gemini 4 scored well on industry benchmarks, internal Google employees reported it performed less well in actual work applications.
Notable quotes
- [02:00] "Now remember, I haven't tested the model, so this is just purely based on what is being currently reported..."
- [07:14] "So that means a 1 million token output limit is not a context window, it's an output limit, meaning that it can write up to a million tokens in a single response."
- [09:46] "...while Gemini 4 has performed well on benchmarks the industry uses to gauge model efficacy, it does less well when employees actually put it to work."
Assessment
This is a secondary commentary and news review video breaking down publicly posted benchmark slides, leaderboard captures, and press reports rather than demonstrating live prompt runs. The presenter explicitly notes that he has not yet personally tested the model and offers balanced skepticism regarding whether the published benchmark scores reflect real-world user workflow efficacy.
Described by gemini-3.8-flash on 2026-10-01 from the video's audio and frames.