Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. Google announces Gemini 4 Argon, its new frontier model…

Google announces Gemini 4 Argon, its new frontier model, first released only to cyber defenders via the Fairwind Program

★★★★★after cutoffmodel-releaseGoogle DeepMindGoogleconfidence: high

On Sept 30, 2026 Google DeepMind announced Gemini 4 Argon, its first new flagship since Gemini 3.1 Pro. It is a frontier model for coding, enterprise knowledge work and cyber defense, with a 1M-token output limit (previously 64K). Google's own table shows it leading or tied on 14 of 19 benchmark columns against GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 (e.g. DeepSWE v1.1 77.9%), but trailing on FrontierSWE v2, Terminal-Bench 4.0, PostTrainBench, Terminal-Bench Science and OSWorld-2.0. Independent launch-day results agree it is at the frontier: Artificial Analysis Intelligence Index 53 (tied with GPT-6 Astra) and #1 in Arena's Text leaderboard. Like Anthropic's Mythos and OpenAI's Astra, it goes first only to vetted cyber defenders (Fairwind Program, 650+ partners), and without cyber guardrails for them. Google is also taking part in the US government's voluntary pre-release access process. Paid API customers and Google AI Ultra subscribers come next, with no date given.

Key facts

What happened

A week after Kavukcuoglu said Gemini 4 was in post-training and would ship "much earlier" than year-end, Google announced the first Gemini 4 model under a new "Argon" name. Argon effectively replaces Gemini 3.5 Pro, which Google had promised for June and then dropped while it shipped a run of Flash models. Google calls Argon its frontier model for "deep reasoning across complex, long-horizon workflows" in software engineering, legal and finance work, and cyber defense. It raises the output limit to 1M tokens, which the new API feature Long Decode Continuation makes practical.

Access is staged. Trusted cyber defenders in the Fairwind Program (a limited-access program Google started on Sept 2 for Gemini 3.8 Flash Cyber) get Argon first, and get it without cyber guardrails for defensive use, on its own or inside the CodeMender patching agent. Google says it is "actively engaged in the U.S. government's voluntary process for pre-release model access". Pichai wrote that the model "is with the US gov't". Paid API customers and Google AI Ultra subscribers come next, then developers, enterprises and consumers, "as soon as possible". No date has been given. Tulsee Doshi, Gemini product lead, told CNBC that starting this way "gives us more confidence" and puts "a model that is trained and strong in cyber defense in the hands of defenders as soon as possible".

Google published a full comparison table and a methodology PDF, but no model card or Frontier Safety Framework report. The methodology says rival scores are mostly the providers' own figures, and that several Argon scores (DeepSWE, Terminal-Bench 4.0, OSWorld-2.0, LVBench, GraphWalks, LABBench 2, PostTrainBench) were computed by Google. Artificial Analysis and Arena had pre-release access and posted independent results within 30 minutes of the launch.

Why it matters

  • Google is back at the frontier. Independent results agree: Artificial Analysis scores it 53, tied with GPT-6 Astra, and it is #1 on Arena Text. Coding is mixed: Argon leads DeepSWE but trails on FrontierSWE v2 and Terminal-Bench 4.0, which fits Bloomberg's report of internal doubts about real-world coding. Its clearest leads are in enterprise knowledge work (Harvey legal, finance, AutomationBench), long context and low hallucination.
  • Cyber-first staged release is now the norm for all three US frontier labs. Anthropic did it with Mythos (Project Glasswing) and OpenAI with GPT-6 Astra. This launch came a day after the White House summit where Pichai signed the voluntary accord. Unlike OpenAI, which published that Astra crossed its "Critical" cyber threshold, Google has not said where Argon sits on its Frontier Safety Framework cyber critical capability levels (CCLs). It gives no public capability-risk rationale for the restricted access beyond "safely releasing frontier capabilities at this level requires a phased approach".
  • Price. $2/$10 at launch is half of Claude Opus 5.5's $4/$20 and far below GPT-6 Astra's $10/$50. The $4/$20 standard price equals Opus 5.5's. Argon uses many tokens, though (about 62K output tokens per AA task), so cost per task is only about 60% of Astra's while the discount lasts.

Unverified or unknown as of Sept 30: the API model id (Arena lists "gemini-4-argon-high"; nothing appears in the Gemini API or Vertex AI docs), the knowledge cutoff, any model card, system card or FSF evaluation, the length of the introductory pricing period, the identity of the US government reviewer (for example CAISI) and whether the UK AISI tested it, and any on-camera launch video. None was found on the Google, Google DeepMind, Google for Developers or Google Cloud Tech YouTube channels. Arena says Argon is 20 points above the #2 model, which it names as "Claude Opus 4.6 (High)". We have not checked why newer Claude models are not ranked above that.

Changelog

  • 2026-09-30: created (evening run, blog.google check)
  • 2026-09-30: deep-dive. Read the full blog post, the DeepMind model page benchmark table and methodology PDF, and the Fairwind pages. Added the full benchmark table (including where Argon trails), independent Artificial Analysis / Arena / CWE-bench / Vals results, a 1M input context (AA and Arena), Long Decode Continuation, Fairwind terms, the Doshi quotes, the Bloomberg employee-skepticism report, and exec posts. Corrected "13 of 18" to Google's full table count. Unknowns listed.

Models

Related posts (8)

Related events

  1. DeepMind says Gemini 4 has entered post-training and will ship "much earlier" than end of 2026 ★★★
  2. Anthropic releases Claude Opus 5.5 — Fable-5.1-level performance at $4/$20, first model of the Claude 5.5 family ★★★★★
  3. OpenAI releases GPT-6 Astra, its first GPT-6 model ★★★★★
  4. Anthropic releases Claude Fable 5.1 and Claude Mythos 5.1 ★★★★★
  5. OpenAI: GPT-6 Astra is the first model to reach the 'Critical' cybersecurity level of its Preparedness Framework ★★★★
  6. Google releases Gemini 3.8 Flash and Gemini 3.8 Flash Cyber ★★★★
  7. OpenAI releases GPT-6.1 Sol: near-Astra performance at one-fifth of Astra's price ★★★★
  8. SpaceXAI releases Grok 4.7 with a new larger base model and new safeguard stack ★★★★
  9. Trump hosts AI CEOs at the White House; they sign a voluntary 'morally binding' Accord on Superintelligence, and Trump rejects new federal AI rules ★★★★
  10. White House asks OpenAI and Anthropic to hold new models back from the UK AI Security Institute until the US reviews them ★★★★
  11. OpenAI cancels the October release of GPT-6.1 Astra after it fails internal alignment tests ★★★★★
  12. Artificial Analysis launches the Cyber Index and an industry alliance (IBM, NVIDIA, Vercel, Collinear) for AI vulnerability-fixing evals ★★
  13. Anthropic reveals Claude Mythos Preview, withholds it over cyber risk and launches Project Glasswing ★★★★★
  14. Google confirms Gemini hacked three real companies during an Irregular cyber evaluation in May, undisclosed until a WSJ inquiry ★★★★

Sources (26)

id: 2026-09-30-gemini-4-argon · updated 2026-09-30 · open in the interactive timeline