Post-Cutoff.com
  1. Home
  2. Posts
  3. Artificial Analysis: Gemini 4 Argon ties GPT-6 Astra (53)…

Artificial Analysis: Gemini 4 Argon ties GPT-6 Astra (53) on the Intelligence Index

Artificial Analysis @ArtificialAnlys · x · 2026-09-30 · ★★★★ · archived

Open the original ↗

First independent evaluation of Argon: Intelligence Index 53 (tie with GPT-6 Astra), lowest hallucination rate among leading models, cost per task, 1M context and the new 'Long Decode Continuation' API feature.

Summary

Artificial Analysis (with pre-release access) reports Gemini 4 Argon (high) at 53 on its Intelligence Index v4.3.2, equal to GPT-6 Astra (max) and one point ahead of GPT-6.1 Sol (max); #1 on AutomationBench-AA (77.5%); Terminal-Bench 4 57% (behind Claude Sonnet 5.5, Opus 5.5 and GPT-6 Astra); AA-Omniscience hallucination rate 15% vs 51% for GPT-6 Astra; averages 62K output tokens per task vs 27K for Astra; $1.99 per Index task at the 50% launch discount ($3.98 at standard price). Lists a 1M-token context window and a new Gemini API feature, Long Decode Continuation, that pauses and resumes long responses across calls to allow up to 1M output tokens without timeouts.

Archived text

Google’s new Gemini 4 Argon equals GPT-6 Astra on the Artificial Analysis Intelligence Index at 60% of the Cost per Task with discounted prices. Google is now back to being one of the top three labs in intelligence achieved

Gemini 4 Argon is @GoogleDeepMind’s first proprietary model above the Flash class in over 7 months. With high reasoning (the highest available), it scores 53 on the Artificial Analysis Intelligence Index, matching GPT-6 Astra (max, 53) and 1 point ahead of GPT-6.1 Sol (max, 52), with gains driven by lower hallucinations and stronger agentic capabilities.

At its current 50% pricing discount and with cache discounts increased to 95%, Gemini 4 Argon costs $1.99 per Intelligence Index task, 60% of GPT-6 Astra (max), but 2.7x GPT-6.1 Sol (max). After the discount ends, this will rise to $3.98 (~1.2x GPT-6 Astra (max)).

Gemini 4 Argon is currently being rolled out to selected users and is not publicly available. The 50% discount is an initial promotion. Google has not yet confirmed the promotion end date

Key benchmarking results for Gemini 4 Argon with high reasoning:

➤ Google returns as one of the top three labs on intelligence: Gemini 4 Argon (high) scores 53 on the Artificial Analysis Intelligence Index, matching GPT-6 Astra (max, 53) and 1 point ahead of GPT-6.1 Sol (max, 52). This is 23 points above Google’s previous non-Flash model, Gemini 3.1 Pro Preview (30) and 12 points ahead of Gemini 3.8 Flash (high)

➤ Launch discounts of 50% make Gemini 4 Argon competitive on Cost per Task: At current discounted pricing, Gemini 4 Argon (high) costs $1.99 per Intelligence Index task, 60% of GPT-6 Astra (max, $3.26) for a comparable level of intelligence. This cost efficiency is driven by lower token prices, rather than reduced token use, with Gemini 4 Argon averaging 62k output tokens per task, compared with 27k for GPT-6 Astra (max). Google has not yet confirmed the promotion end date, but on standard pricing, Cost per Task will increase to $3.98

➤ Stronger agentic performance: Historically a weaker area for Gemini models, Gemini 4 Argon shows improvements across agentic evaluations. It ranks #1 on AutomationBench-AA at 77.5%, 6 points ahead of Claude Sonnet 5.5 (max, 71.3%). On Terminal Bench 4, Gemini 4 Argon achieves 57%, a +53 point improvement from Gemini 3.1 Pro Preview, only behind Claude Sonnet 5.5 (max, 64%), Claude Opus 5.5 (max, 60%) and GPT-6 Astra (59%). On AA-Briefcase, it reaches 1494 Elo. This is driven by a 65% rubric pass rate, the highest we have recorded, but lower Analytical Quality (1576 Elo) and Presentation Quality (1308 Elo)

➤ Lowest hallucination rate among leading models: On AA-Omniscience, Gemini 4 Argon has a 15% hallucination rate, the lowest of any model scoring 45+ on the Intelligence Index, compared with 51% for GPT-6 Astra (max) and 54% for GPT-6.1 Sol (max). This means Argon is much more likely to acknowledge when it does not know an answer rather than guess incorrectly. On accuracy, Gemini 4 Argon scores 50%, a 5 point decrease from Gemini 3.1 Pro Preview, and 13 points below GPT-6 Astra (max, 63%). With this slightly lower accuracy, its overall AA-Omniscience score of 42 remains in line with GPT-6 Astra (43) and GPT-6.1 Sol (42)

Key model details:

➤ Context Window: 1M tokens

➤ Multimodality: Text, image, video, and speech input, with text output

➤ Pricing: $4/$20 per 1M input/output tokens at standard pricing, currently discounted 50% to $2/$10. Cached input tokens receive a 95% discount ($0.10 per 1M at discounted pricing), up from 90% on Gemini 3.8 Flash

➤ Long Decode Continuation: We tested Gemini 4 Argon with Long Decode Continuation, a new Gemini API feature that pauses long responses and resumes them across follow-up calls. This lets reasoning run up to 1M output tokens without request timeouts

Media: https://pbs.twimg.com/media/HTfY3plasAAay9f.jpg?name=orig

views 8856 · likes 273 · reposts 29 · replies 15 (at fetch time)

Archived 2026-09-30 via fxtwitter (unofficial).

Related events

All posts · id: 2026-09-30-artificialanalysis-gemini-4-argon