Gemini 4 Argon | First Impressions
Arena AI · 2026-09-30 · review · 34,565 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Peter Gostev, AI Capability Lead at Arena, gives his first impressions of Google's newly released Gemini 4 (Gemini 4 Argon High). He examines where the model lands on Arena's Agent Arena leaderboard and Pareto frontier, and compares its interactive 3D web generation capabilities against competing models like Claude Sonnet 5.5 and GPT-6.1 Sol.
What is shown
- [00:12] Agent Arena Leaderboard: Overall dynamic ranking table where Gemini 4 Argon (High) is shown ranked 8th.
- [00:33] Pareto Frontier Chart: Plotting net improvement against median cost per task, showing Gemini 4 Argon High positioned on the Pareto frontier.
- [00:56] Golden Gate & SF Bay 3D Scene: Gemini 4 Argon High’s interactive 3D rendering with weather/traffic controls, compared against Claude Sonnet 5.5 [01:23] and GPT-6.1-Sol Max [01:38].
- [02:12] Floating Islands ("Sanctuary of the Sky Cataracts"): Gemini 4 Argon High's generation showing geometry and tree glitches, contrasted with Claude Sonnet 5.5 [02:33] and GPT-6.1-Sol Max [02:56].
- [03:14] Paris Flight Simulator Game: Gemini 4 Argon High generating an Eiffel Tower airplane ring game with clunky controls, compared with Claude Sonnet 5.5 [03:39] and GPT-6.1-Sol Max [03:48].
- [04:06] Bruegel's Tower of Babel: Gemini 4 Argon High's 3D generation showing floating, blocky character sprites, compared with GPT-6.1-Sol Max [04:52].
- [05:16] The White House: Gemini 4 Argon High scene displaying visual artifacts on the roof and landscaping, compared with Claude Sonnet 5.5 [05:51] and GPT-6.1-Sol Max [06:12].
- [06:37] Sagrada Família: Gemini 4 Argon High 3D scene showing structural disconnects and geometry issues.
- [08:00] Prompt Vault & Cost Analysis: Reviewing Prompt Vault comparative tables and noting Gemini 4 Argon's price-performance ratio on the Pareto curve [08:02].
Claims & numbers
- The presenter notes that Gemini 4 Argon (High) ranks 8th overall on the Agent Arena leaderboard, placing just above Sol 5/6 and below models like Claude Fable 5.1, Claude Opus 5.5, GPT-6 Astra, GPT-6 Sol, Claude Fable 5, and Claude Opus 5 [00:17].
- The presenter notes that Gemini 4 Argon makes the Pareto-optimal frontier [00:40].
- The presenter states that Gemini 4 Argon High costs roughly one-third less per task than GPT-6 Sol (~$0.35 vs ~$0.92 on the displayed chart) and roughly 70% cheaper than Claude Opus 5.5 [08:18].
- The presenter argues that while Gemini 4 Argon is cost-effective, its quality, stability, and aesthetic coherence in complex 3D artifact generation lag behind Claude Sonnet 5.5 and GPT-6.1 Sol [08:30].
Notable quotes
- "So it ranks eighth, so it's just above Sol 5.6, but it is below some of the latest models such as for example, 6 Sol or Fable 5 or Opus 5." [00:17]
- "Gemini models tend to be quite efficient in terms of, or at least well-priced, and it does make the Pareto frontier, so this is important, but it's kind of in the middle of the pack." [00:35]
- "In terms of the performance, it is lower than others, and these kinds of examples demonstrate that pretty well." [08:30]
Assessment
This is a genuine third-party review and analysis video showing interactive 3D browser environments generated from identical prompts across multiple frontier models. The demonstrations are authentic screen recordings highlighting genuine strengths, bugs, and control issues without obvious deceptive editing.
Described by gemini-3.8-flash on 2026-10-01 from the video's audio and frames.