Google announces Gemini 4 Argon, its new frontier model, first released only to cyber defenders via the Fairwind Program
On Sept 30, 2026 Google DeepMind announced Gemini 4 Argon, its first new flagship since Gemini 3.1 Pro. It is a frontier model for coding, enterprise knowledge work and cyber defense, with a 1M-token output limit (previously 64K). Google's own table shows it leading or tied on 14 of 19 benchmark columns against GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 (e.g. DeepSWE v1.1 77.9%), but trailing on FrontierSWE v2, Terminal-Bench 4.0, PostTrainBench, Terminal-Bench Science and OSWorld-2.0. Independent launch-day results agree it is at the frontier: Artificial Analysis Intelligence Index 53 (tied with GPT-6 Astra) and #1 in Arena's Text leaderboard. Like Anthropic's Mythos and OpenAI's Astra, it goes first only to vetted cyber defenders (Fairwind Program, 650+ partners), and without cyber guardrails for them. Google is also taking part in the US government's voluntary pre-release access process. Paid API customers and Google AI Ultra subscribers come next, with no date given.
Key facts
- Announced Sept 30, 2026 (~20:00 UTC) by Koray Kavukcuoglu (SVP, Google DeepMind and Chief AI Architect) on the Google blog; Sundar Pichai called it 'an early look' given 'lots of discussion out there about our next model(!)'
- 'Argon' replaces the promised Gemini 3.5 Pro, which Google had announced at I/O in May for June but never shipped (Ars Technica, The New Stack, 9to5Google)
- Output limit: 1M tokens, up from 64K (Google). Input context window: 1M tokens per Artificial Analysis and Arena's leaderboard metadata. Google's blog does not state it
- New Gemini API feature 'Long Decode Continuation' pauses long responses and resumes them across follow-up calls, allowing up to 1M output tokens without request timeouts (Artificial Analysis, which tested it)
- Price: $2 / $10 per 1M input/output tokens at launch (a 50% introductory discount), then $4 / $20; cached input 95% off, i.e. $0.10 per 1M at the launch price (Google; Artificial Analysis). No end date for the discount
- Vendor-reported benchmarks (Google table; rivals' numbers mostly their self-reported figures): Vals Index 68.9% (GPT-6 Astra 63.1, Fable 5.1 65.8, Opus 5.5 67.0); AutomationBench 51.3% (Opus 5.5 42.5); Vals Finance Agent v2 65.4%; Harvey Legal Agent Benchmark 19.6% (next best 6.7); DeepSWE v1.1 77.9% (Opus 5.5 74.2, Astra 74.1); Vibe Code Bench 91.9%; LABBench 2 88.8%; RiemannBench 76.0%; GraphWalks 256K–1M 84.2% (Astra 71.8); Agent's Last Exam 39.5%; Chartography 71.6%; LVBench 91.7%; CWE-bench v1 68.0% (tied with Astra)
- Where Argon trails (Google's own table): FrontierSWE v2 55.0% (Astra 65.5), Terminal-Bench 4.0 57.4% (Opus 5.5 66.4), PostTrainBench 45.3% (Opus 5.5 49.3), Terminal-Bench Science 0.1 57.6% (Astra 68.1), OSWorld-2.0 offline partial score 69.2% (Astra 72.6)
- Cyber (Google): 85.8% on Google's internal vulnerability-discovery benchmark and 70.9% on Wiz's black-box penetration-testing benchmark, vs 71.0% and 58.2% for Gemini 3.8 Flash Cyber (figures via The New Stack); no rival numbers given. Wiz 'Scan for Good' used it to find a critical flaw in hospital healthcare software that 'previous frontier models had missed' (no specifics)
- Independent: Artificial Analysis Intelligence Index 53 (high reasoning), equal to GPT-6 Astra (max) and 1 point above GPT-6.1 Sol; 15% hallucination rate on AA-Omniscience (Astra 51%); ~62K output tokens per task (Astra 27K); $1.99 per Index task at the launch price, $3.98 at the standard price
- Independent: Arena Text leaderboard #1 at 1525 (±~9; 4,942 votes; listed as pre-release), 20 points above #2; #8 in Code Arena WebDev (1679)
- Independent leaderboards: CWE-bench v1 68% pass@1 (75% pass@4) in the Antigravity harness, tied with Grok 4.7 and GPT-6 Astra but at $6.63 per rollout vs $0.79 for Claude Opus 5.5 (67%); Vals AI lists Vals Index 68.9%
- Initial access: Fairwind Program only (launched Sept 2 with Gemini 3.8 Flash Cyber; 650+ partners: governments and national cyber authorities, critical-infrastructure operators, core tech platforms). Partners must use user-level authentication and phishing-resistant MFA, limit access to internal security, incident-response or pentest teams, and may not resell access. Zero data retention is available when Argon is used as a managed model
- Internal use (Google): thousands of Googlers use it, including in Antigravity. Argon agents freed 300+ TiB of data-center memory (500 TiB–1 PiB expected), are migrating C/C++ to Rust (re2, libgav1, 800K+ lines of the Fuchsia Zircon kernel), made a libgav1 Rust port 2.7x faster by replacing 32K lines of SIMD code, and beat a published quantum-algorithm spacetime-cost baseline by 40%
- Safety (Google): CBRN and cyber misuse refusals under the Frontier Safety Framework; activation-based misuse monitoring; internal and external red teams; 'most resilient model yet' against indirect prompt injection (leads Gray Swan IPI); chain-of-thought and action monitors that can stop execution, with training-run monitoring kept out of the training signal; sandboxes sealed before high-risk training or evals. No model card, FSF critical-capability-level report or system card published at launch
- Bloomberg (Sept 30): some Google employees say it does well on benchmarks but less well in real work and 'struggles to handle certain coding tasks'; Google called that characterization inaccurate, and one employee cited a 'large consensus' that it is at the frontier
- Before launch: codenamed 'argon'; mid-September leaks described a 256K output limit (the final figure is 1M)
What happened
A week after Kavukcuoglu said Gemini 4 was in post-training and would ship "much earlier" than year-end, Google announced the first Gemini 4 model under a new "Argon" name. Argon effectively replaces Gemini 3.5 Pro, which Google had promised for June and then dropped while it shipped a run of Flash models. Google calls Argon its frontier model for "deep reasoning across complex, long-horizon workflows" in software engineering, legal and finance work, and cyber defense. It raises the output limit to 1M tokens, which the new API feature Long Decode Continuation makes practical.
Access is staged. Trusted cyber defenders in the Fairwind Program (a limited-access program Google started on Sept 2 for Gemini 3.8 Flash Cyber) get Argon first, and get it without cyber guardrails for defensive use, on its own or inside the CodeMender patching agent. Google says it is "actively engaged in the U.S. government's voluntary process for pre-release model access". Pichai wrote that the model "is with the US gov't". Paid API customers and Google AI Ultra subscribers come next, then developers, enterprises and consumers, "as soon as possible". No date has been given. Tulsee Doshi, Gemini product lead, told CNBC that starting this way "gives us more confidence" and puts "a model that is trained and strong in cyber defense in the hands of defenders as soon as possible".
Google published a full comparison table and a methodology PDF, but no model card or Frontier Safety Framework report. The methodology says rival scores are mostly the providers' own figures, and that several Argon scores (DeepSWE, Terminal-Bench 4.0, OSWorld-2.0, LVBench, GraphWalks, LABBench 2, PostTrainBench) were computed by Google. Artificial Analysis and Arena had pre-release access and posted independent results within 30 minutes of the launch.
Why it matters
- Google is back at the frontier. Independent results agree: Artificial Analysis scores it 53, tied with GPT-6 Astra, and it is #1 on Arena Text. Coding is mixed: Argon leads DeepSWE but trails on FrontierSWE v2 and Terminal-Bench 4.0, which fits Bloomberg's report of internal doubts about real-world coding. Its clearest leads are in enterprise knowledge work (Harvey legal, finance, AutomationBench), long context and low hallucination.
- Cyber-first staged release is now the norm for all three US frontier labs. Anthropic did it with Mythos (Project Glasswing) and OpenAI with GPT-6 Astra. This launch came a day after the White House summit where Pichai signed the voluntary accord. Unlike OpenAI, which published that Astra crossed its "Critical" cyber threshold, Google has not said where Argon sits on its Frontier Safety Framework cyber critical capability levels (CCLs). It gives no public capability-risk rationale for the restricted access beyond "safely releasing frontier capabilities at this level requires a phased approach".
- Price. $2/$10 at launch is half of Claude Opus 5.5's $4/$20 and far below GPT-6 Astra's $10/$50. The $4/$20 standard price equals Opus 5.5's. Argon uses many tokens, though (about 62K output tokens per AA task), so cost per task is only about 60% of Astra's while the discount lasts.
Unverified or unknown as of Sept 30: the API model id (Arena lists "gemini-4-argon-high"; nothing appears in the Gemini API or Vertex AI docs), the knowledge cutoff, any model card, system card or FSF evaluation, the length of the introductory pricing period, the identity of the US government reviewer (for example CAISI) and whether the UK AISI tested it, and any on-camera launch video. None was found on the Google, Google DeepMind, Google for Developers or Google Cloud Tech YouTube channels. Arena says Argon is 20 points above the #2 model, which it names as "Claude Opus 4.6 (High)". We have not checked why newer Claude models are not ranked above that.
Changelog
- 2026-09-30: created (evening run, blog.google check)
- 2026-09-30: deep-dive. Read the full blog post, the DeepMind model page benchmark table and methodology PDF, and the Fairwind pages. Added the full benchmark table (including where Argon trails), independent Artificial Analysis / Arena / CWE-bench / Vals results, a 1M input context (AA and Arena), Long Decode Continuation, Fairwind terms, the Doshi quotes, the Bloomberg employee-skepticism report, and exec posts. Corrected "13 of 18" to Google's full table count. Unknowns listed.
Models
- Gemini 4 Argon Google DeepMind · preview
Related posts (8)
- Artificial Analysis: Gemini 4 Argon ties GPT-6 Astra (53) on the Intelligence Index original ↗ Artificial Analysis @ArtificialAnlys · x · 2026-09-30
First independent evaluation of Argon: Intelligence Index 53 (tie with GPT-6 Astra), lowest hallucination rate among leading models, cost per task, 1M context and the new 'Long Decode Continuation' API feature. - Google DeepMind: Introducing Gemini 4 Argon original ↗ Google DeepMind @GoogleDeepMind · x · 2026-09-30
Official lab launch post, the most-viewed launch post (~254K views at fetch). - Sundar Pichai introduces Gemini 4 Argon with a benchmark table original ↗ Sundar Pichai @sundarpichai · x · 2026-09-30
CEO announcement with the widest reach of any Google exec post on the launch (200K+ views); 'Lots of discussion out there about our next model(!)' frames it as an early look ahead of broad release. - Arena: Gemini 4 Argon (High) debuts #1 in Text Arena (1525) original ↗ Arena.ai @arena · x · 2026-09-30
Independent crowd-preference leaderboard result on launch day. - Logan Kilpatrick: Gemini 4 Argon priced at $2 in / $10 out during introductory pricing original ↗ Logan Kilpatrick @OfficialLoganK · x · 2026-09-30
Gemini API product lead confirming the introductory API price to developers (~147K views). - Sundar Pichai: Argon is with the US government and rolling out responsibly original ↗ Sundar Pichai @sundarpichai · x · 2026-09-30
First-hand CEO statement that the model was given to the US government before release and that access is staged for safety reasons. - Koray Kavukcuoglu: sharing Argon 'as soon as possible', starting with Fairwind defenders original ↗ koray kavukcuoglu @koraykv · x · 2026-09-30
Statement by the executive who authored the launch blog and runs Google DeepMind day to day (low reach, ~2.5K views). - Lentils original ↗ Lentils @Lentils80 · x · 2026-09-14
Cited as a source by: 2026-09-30-gemini-4-argon
Related events
- DeepMind says Gemini 4 has entered post-training and will ship "much earlier" than end of 2026 ★★★
- Anthropic releases Claude Opus 5.5 — Fable-5.1-level performance at $4/$20, first model of the Claude 5.5 family ★★★★★
- OpenAI releases GPT-6 Astra, its first GPT-6 model ★★★★★
- Anthropic releases Claude Fable 5.1 and Claude Mythos 5.1 ★★★★★
- OpenAI: GPT-6 Astra is the first model to reach the 'Critical' cybersecurity level of its Preparedness Framework ★★★★
- Google releases Gemini 3.8 Flash and Gemini 3.8 Flash Cyber ★★★★
- OpenAI releases GPT-6.1 Sol: near-Astra performance at one-fifth of Astra's price ★★★★
- SpaceXAI releases Grok 4.7 with a new larger base model and new safeguard stack ★★★★
- Trump hosts AI CEOs at the White House; they sign a voluntary 'morally binding' Accord on Superintelligence, and Trump rejects new federal AI rules ★★★★
- White House asks OpenAI and Anthropic to hold new models back from the UK AI Security Institute until the US reviews them ★★★★
- OpenAI cancels the October release of GPT-6.1 Astra after it fails internal alignment tests ★★★★★
- Artificial Analysis launches the Cyber Index and an industry alliance (IBM, NVIDIA, Vercel, Collinear) for AI vulnerability-fixing evals ★★
- Anthropic reveals Claude Mythos Preview, withholds it over cyber risk and launches Project Glasswing ★★★★★
- Google confirms Gemini hacked three real companies during an Irregular cyber evaluation in May, undisclosed until a WSJ inquiry ★★★★
Sources (26)
- officialGoogle: Gemini 4 Argon, our next era of frontier intelligence (Koray Kavukcuoglu)
- officialGoogle DeepMind: Gemini model page (full benchmark table)
- officialGoogle DeepMind: Gemini 4 Argon model evaluation, approach & methodology (PDF)
- officialGoogle DeepMind: Fairwind Program
- officialGoogle: Proactive cyber defense for governments and enterprises (Fairwind launch, Sept 2)
- officialDeepMind Institute: The case for reasoning transparency (linked from the launch post)
- discussionSundar Pichai on X: Introducing Gemini 4 Argon
- discussionArtificial Analysis on X: Gemini 4 Argon evaluation
- docsArtificial Analysis: Gemini 4 Argon model page
- discussionArena on X: Gemini 4 Argon #1 in Text Arena
- docsArena Text leaderboard
- docsCWE-bench leaderboard
- docsVals AI: Vals Index
- pressCNBC: Google rolls out Gemini 4 Argon, its most advanced AI model
- pressAxios: Google unveils Gemini 4, long-awaited answer to OpenAI and Anthropic
- pressBloomberg: Google grapples with employee skepticism about new Gemini model
- pressReuters: Google announces Gemini 4 flagship AI model after months of delays
- pressNYT: Google releases a new flagship AI model, with limits
- pressArs Technica: Google announces Gemini 4 Argon AI model, but you can't use it yet
- pressThe New Stack: Gemini 4 Argon is here, it's great, and you can't have it yet
- pressThe Next Web: Gemini 4 Argon, Google's new flagship reaches cyber defenders first
- press9to5Google: Google announces Gemini 4 Argon as its new frontier model
- pressTestingCatalog: Google unveils Gemini 4 Argon with SOTA score on DeepSWE
- pressVentureBeat: Google unveils Gemini 4 Argon, retaking benchmark lead, but in limited release
- discussionHacker News discussion
- discussionLentils on X: first 'argon' output leak (Sept 2026)
id: 2026-09-30-gemini-4-argon · updated 2026-09-30 · open in the interactive timeline