Build with AI
Gemini 4 Argon: Everything We Know About Google's New Frontier Model (2026)
Google announced Gemini 4 Argon on September 30, 2026, introducing a 1M output token window, a 77.9% state of the art on DeepSWE v1.1, and an introductory API price of $2 per million input tokens and $10 per million output tokens. Access is currently restricted to vetted cyber defenders through the Fairwind Program, with paid API and Ultra access scheduled next. Here is the full breakdown of benchmarks, real-world Google production data, and rollout timelines.

Contents (13)
- Core specifications and launch facts
- How Google arrived at Gemini 4 Argon
- The 1M output token window and long-horizon reasoning
- Independent benchmarks: Strengths and gaps
- Internal production proof points at Google
- Defensive cybersecurity and the Fairwind rollout
- Safety mechanisms and government oversight
- API pricing and token economics
- Pre-launch community leaks versus official reality
- Access timeline: When can you actually use it?
- Navigating model changes as an AI software founder
- Frequently asked questions
- Primary sources and reference documentation
Last updated: September 30, 2026. This story reflects official launch documentation from Google DeepMind and independent benchmark evaluations published on announcement day.
Google announced Gemini 4 Argon on September 30, 2026, marking the company's return to the frontier Pro class after ten months of Flash-tier releases. The model serves as the flagship release for the Gemini 4 generation, carrying an official tagline of "our next era of frontier intelligence."
Argon introduces an industry-leading output ceiling of 1,000,000 tokens, a state-of-the-art score of 77.9% on DeepSWE v1.1, and an introductory API rate of $2.00 per million input tokens and $10.00 per million output tokens. General developers and consumers cannot use it yet. Google restricted initial deployment to vetted defense organizations, critical infrastructure teams, and select government partners through its gated Fairwind Program. Access for Google AI Ultra subscribers and paid API customers is scheduled to follow as soon as guardrail evaluations finish.
A note on perspective: SaaSCity tracks the products, infrastructure, and economics of independent software companies. Hundreds of software founders listed on our directory build agents and workflows on top of external frontier models. When an upstream provider changes output limits, cuts token rates, or delays public access, those decisions ripple through software unit margins. If you are building a product, you can list it on the SaaSCity directory for free to claim a building on our live map. Founders needing rapid verification can choose Quick Pass for 24-hour review ($19.99), or Premium for a featured launch post and permanent backlinks ($99.99).
Here is the verified record of what Gemini 4 Argon delivers, where it leads, where independent benchmarks show gaps, and when wider access will open.
Core specifications and launch facts
The table below summarizes the verified technical specifications, pricing tiers, and operational parameters confirmed by Google DeepMind on September 30, 2026.
| Parameter | Confirmed specification | Source / Notes |
|---|---|---|
| Model Name | Gemini 4 Argon | First model of the Gemini 4 generation |
| Developer | Google DeepMind | Primary announcement by Koray Kavukcuoglu |
| Predecessor Path | Gemini 3 Pro (Nov 2025) | Gemini 3.5 Pro skipped after summer delays |
| Model Class | Frontier / Pro flagship | Larger parameter scale than prior Pro models |
| Output Token Limit | 1,000,000 tokens | 16x increase over previous 64K Gemini limit |
| Input Context Window | 1M-token class | Frontier-standard context window |
| Introductory API Input | $2.00 / 1M tokens | Standard rate before cache discounts |
| Introductory API Output | $10.00 / 1M tokens | High reasoning output |
| Prompt Caching Discount | 95% off input | $0.10 / 1M cached input tokens |
| Standard API Input | $4.00 / 1M tokens | Effective following introductory window |
| Standard API Output | $20.00 / 1M tokens | Matches Claude Opus 5.5 output baseline |
| Current Availability | Fairwind Program only | Gated to vetted cyber defenders and governments |
| Next Availability Wave | Paid API & Google AI Ultra | Stated as "as soon as possible" |
| Consumer Availability | Not available | No public web interface date set |

How Google arrived at Gemini 4 Argon
Google's frontier release cadence experienced an unusual pause during 2026. The company last updated its top-tier weights in November 2025 with Gemini 3 Pro. Google had internally targeted June 2026 for Gemini 3.5 Pro, but performance benchmarks failed to meet internal targets. Google delayed the checkpoint, reworked the underlying training runs, and ultimately skipped the 3.5 version number entirely.
Between December 2025 and late September 2026, Google relied on iterative Flash-class releases to stay visible. The most notable checkpoint was Gemini 3.8 Flash Cyber, released on September 2, 2026 alongside CodeMender to anchor the Fairwind Program. While Flash models delivered low latency and competitive pricing for lightweight tasks, developers building complex agents frequently opted for Anthropic's Claude Opus 5.5 or OpenAI's GPT-6 Astra for long-horizon reasoning.
Google spokespeople confirmed that Gemini 4 Argon is physically larger than prior Pro-class systems. The release also inaugurates an elemental naming convention. Google adopted elemental designations like Argon to distinguish flagship architectures from its lighter Flash lines.
The 1M output token window and long-horizon reasoning
Frontier models have offered million-token input windows for over a year, but generation output has remained constrained. Prior Gemini checkpoints capped completions at 64,000 tokens, and competing frontier models typically stop between 32,000 and 128,000 tokens per request.
Gemini 4 Argon raises the single-turn output ceiling to 1,000,000 tokens. In his launch post, Koray Kavukcuoglu explained the operational rationale:
"When the model has the headroom to think deeply and generate hundreds of thousands of tokens in a single trajectory, it adds a new level of depth in reasoning to solve tough problems in one go."
For software engineering agents, an expanded generation ceiling eliminates the need to break complex tasks into fragile intermediate conversational turns. An agent can ingest a multi-file dependency graph, produce hundreds of thousands of tokens of internal chain-of-thought analysis, synthesize new modules, and output complete files in a single continuous stream without running into buffer boundaries.
Independent benchmarks: Strengths and gaps
Google centered its announcement on two headline claims: setting a new state of the art on DeepSWE v1.1 with 77.9%, and taking the top position on the Vals Index with 68.9%. Independent benchmark runs from Vals AI, Artificial Analysis, and The New Stack confirm strong results in enterprise knowledge work and specific coding tasks, while also identifying benchmarks where Argon trails GPT-6 Astra and Claude Opus 5.5.
The table below compiles published evaluations as of September 30, 2026.
| Category | Benchmark | Gemini 4 Argon | GPT-6 Astra | Claude Opus 5.5 | Claude Fable 5.1 |
|---|---|---|---|---|---|
| Agentic coding | DeepSWE v1.1 | 77.9% (SOTA) | 74.1% | 74.2% | 67.4% |
| Knowledge work | Vals Index | 68.9% (#1) | 63.1% | 67.0% | 65.8% |
| Knowledge work | AutomationBench | 51.3% (#1) | 41.4% | 42.5% | 31.4% |
| Finance | Vals Finance Agent v2 | 65.4% | 53.5% | 58.6% | 58.9% |
| Legal | Harvey Legal Agent | 19.6% | 5.4% | 3.8% | 6.7% |
| Cyber defense | CWE-bench v1 | 68% (Tie) | 68% | 67% | 58% |
| Video understanding | LVBench | 91.7% (SOTA) | 87.5% | 83.7% | 79.7% |
| Agentic coding | FrontierSWE v2 | 55.0% | 65.5% | 62.3% | 56.3% |
| Terminal commands | Terminal-bench 4.0 | 57.4% | 58.2% | 66.4% | 57.9% |
| Long context | GraphWalks 256k-1M BFS F1 | 84.2% | 71.8% | 66.8% | 65.0% |
| Chart reasoning | Chartography | 71.6% | 71.0% | 66.3% | 46.2% |

Where Gemini 4 Argon leads
The model shows its clearest advantage on tasks requiring sustained reasoning across diverse document types and multi-step business logic:
- DeepSWE v1.1 (77.9%): DeepSWE evaluates real-world software engineering across 113 contamination-resistant tasks, 91 public code repositories, and 5 programming languages. Argon's 77.9% pass rate leads GPT-6 Astra (74.1%) by 3.8 percentage points and Claude Opus 5.5 (74.2%) by 3.7 points.
- Enterprise knowledge work: On the Vals Index, which weights economic value across finance, legal, tax, and coding by each sector's contribution to U.S. GDP, Argon captured first place with 68.9%. Vals AI reported that Argon placed in the top five on 20 out of 22 evaluated benchmarks.
- Specialized vertical agents: On the Harvey Legal Agent benchmark, Argon scored 19.6%, outperforming GPT-6 Astra (5.4%) and Claude Opus 5.5 (3.8%). On Vals Finance Agent v2, Argon reached 65.4%, compared to 53.5% for Astra.
- Multimodal video and documents: On LVBench, which measures long video comprehension and fine-grained temporal retrieval, Argon reached 91.7%, ahead of Astra's 87.5%.
Where Gemini 4 Argon trails
Argon does not hold an across-the-board sweep. Two established benchmarks highlight areas where competing models retain advantages:
- FrontierSWE v2: On FrontierSWE v2, another demanding agentic coding evaluation, Argon achieved 55.0%. GPT-6 Astra leads this benchmark with 65.5%, and Claude Opus 5.5 reached 62.3%.
- Terminal-bench 4.0: Testing command-line execution and shell problem-solving, Claude Opus 5.5 leads comfortably at 66.4%. Argon registered 57.4%, trailing both Opus 5.5 and Astra (58.2%).
On Artificial Analysis's Intelligence Index, Gemini 4 Argon scored 53 points under maximum reasoning effort. That score ties GPT-6 Astra (53) and sits one point above GPT-6.1 Sol (52). Artificial Analysis also noted that Argon ranked first on its AutomationBench-AA suite with 77.5%, re-establishing Google in the top-three frontier tier for the first time in over seven months.
Internal production proof points at Google
Independent benchmarks test synthetic or historical problems, but Google provided five production deployments where internal engineering teams put Argon to work on live infrastructure before launch.

1. Data-center memory reclamation
Teams deployed an ensemble of Argon agents to analyze fleet-wide profiling telemetry across Google data centers. The agents identified redundant memory allocations and applied autonomous software patches, freeing over 300 TiB of physical RAM across active clusters. Google engineering projects these optimizations will reclaim between 500 TiB and 1 PiB of memory once rolled across all global server pools.
2. Large-scale C/C++ to Rust migrations
Memory safety remains a central operational priority for cloud providers. Google used Argon to convert legacy C and C++ codebases into idiomatic Rust. Conversions completed to date include tens of thousands of lines in the re2 regular expression library and the libgav1 media decoder. The agents also generated conversion drafts for over 800,000 lines of the Fuchsia OS Zircon kernel, which are undergoing formal audit and emulation checks.
3. SIMD video decoding performance
During the libgav1 video decoder migration, Argon replaced approximately 32,000 lines of SIMD assembly in an existing Rust port. The resulting implementation executed 2.7 times faster than the previous Rust code while producing bit-identical output, matching the performance curve of handwritten, hand-tuned C++.
4. Quantum algorithm compilation
Google's quantum computing division used Argon to optimize the spacetime volume (physical qubits multiplied by gate cycles) for critical subroutines. The model analyzed the compilation pipelines and produced an allocation schedule that beat published academic baselines by 40% in minutes of execution.
5. Critical infrastructure vulnerability discovery
Through a collaboration with cloud security firm Wiz, an un-sandboxed defense instance of Argon scanned healthcare software used in hospital networks worldwide. The model discovered a high-severity data exposure vulnerability affecting sensitive medical records that previous frontier reasoning models had passed over.
Defensive cybersecurity and the Fairwind rollout
The central public policy dimension of Gemini 4 Argon is its gated release. Google restricted initial availability to trusted defensive organizations through the Fairwind Program.

Google established the Fairwind Program on September 2, 2026, pairing Gemini 3.8 Flash Cyber with the CodeMender automated remediation tool. Argon is the second and most capable model added to that program.
Access within Fairwind is limited to:
- National cybersecurity authorities and computer emergency response teams
- Vetted government agencies engaged in infrastructure protection
- Critical utility, transportation, and healthcare network operators
- Trusted security research partners like Wiz under the Scan for Good initiative
Google's stated reason for gating the model is capability duality. Software engineering models capable of finding and remediating obscure buffer overflows, logical flaws, and zero-day vulnerabilities can also be prompted to generate working attack exploits. To maximize defensive utility, Google supplies Fairwind members with a version of Argon that bypasses general cyber guardrails, allowing defensive operators to inspect live exploit surfaces without artificial refusal triggers.
On CWE-bench v1, which evaluates autonomous remediation of Common Weakness Enumeration security flaws, Gemini 4 Argon achieved a 68% success rate, tying GPT-6 Astra for first place.
Safety mechanisms and government oversight
Google DeepMind highlighted several technical safeguards developed to monitor Argon's reasoning chains and prevent misuse before general API availability:
- Activation monitoring: Google monitors the model's internal activations and chain-of-thought traces directly during generation, supplementing standard input filtering. If intermediate reasoning signals intent toward chemical, biological, radiological, nuclear (CBRN), or offensive cyber generation, runtime monitors intervene to abort generation.
- Indirect prompt injection defense: On Gray Swan's Indirect Prompt Injection (IPI) benchmark, Google reports that Argon achieved top resistance against hidden prompt overrides embedded in third-party websites and ingested files.
- Federal pre-release review: Google submitted Gemini 4 Argon to the voluntary pre-release access program coordinated by the U.S. government, providing federal evaluators access to assess national security risks prior to civilian API release.
- Safety accord alignment: The announcement coincided with Sundar Pichai joining other artificial intelligence laboratory executives in signing a voluntary safety framework with federal officials in Washington. The agreement establishes non-binding principles for testing and sharing vulnerability research across frontier developers.
API pricing and token economics
When the API opens to commercial developers, Google plans to introduce Argon with aggressive promotional rates before moving to standard pricing.
The table below breaks down the cost structure per million tokens:
| Token tier | Promotional rate | Standard rate | Claude Opus 5.5 comparison |
|---|---|---|---|
| Input Tokens (Prompt) | $2.00 / 1M | $4.00 / 1M | ~$2.00 / 1M |
| Output Tokens (Generation) | $10.00 / 1M | $20.00 / 1M | ~$20.00 / 1M |
| Cached Input Tokens | $0.10 / 1M (95% off) | Variable | ~$0.50 / 1M |
During the introductory period, Argon offers a notable price advantage over frontier peers. Artificial Analysis calculated that for typical tasks on its Intelligence Index, running Argon with prompt caching costs approximately $1.99 per benchmark task. That figure represents roughly 60% of the cost to run the same tasks on GPT-6 Astra max. After promotional pricing concludes and output moves to $20.00 per million, the average task cost will rise to $3.98, putting it roughly in line with Astra and Claude Opus 5.5.
For software founders building autonomous agents, prompt caching provides the real economic wedge. Agent loops that repeatedly pass repository structure, system prompt instructions, and tool definitions can cut input expenses by 95%, reducing base input to $0.10 per million tokens.
Pre-launch community leaks versus official reality
During mid-September 2026, anonymous checkpoints surfaced across community testing platforms like LMSYS Arena and developer chat rooms under handles like "Lentils" and "gemini-3.8-flash-exp." Speculation spread quickly:
- Output token rumors: Early testers observed checkpoints generating up to 256,000 tokens and assumed that was the final limit. Google's official announcement delivered 1,000,000 output tokens.
- Thinking latency: Testers noted reasoning traces lasting over two minutes on complex mathematical problems, matching Google's confirmed high-effort configuration.
- Context window rumors: Unverified forum threads claimed Argon possessed a 2M-token or 10M-token context window. Google's documentation places the model in the standard 1M context class, prioritizing output depth over expanded input length.
- Frontend design: Early leakers highlighted improved frontend styling, cleaner Tailwind CSS output, and competent SVG rendering. While Google's blog did not spotlight CSS generation, third-party testers noted higher visual fidelity on website rendering tasks.
Access timeline: When can you actually use it?
Google has outlined three distinct phases for Gemini 4 Argon distribution:
- Phase 1 (Live today, September 30, 2026): Vetted defense teams, intelligence agencies, critical infrastructure operators, and security partners accessing Argon via the Fairwind Program without general cyber guardrails.
- Phase 2 (Upcoming, date unannounced): Paid Google AI Studio and Vertex AI customers, along with Google AI Ultra consumer subscribers. Google states this rollout will occur "as soon as possible" as engineers iterate on runtime guardrails.
- Phase 3 (Unconfirmed): Free consumer web availability within the standard Gemini chat interface. Google has made no commitments regarding if or when a full-scale Argon model will be accessible to free-tier users.
Navigating model changes as an AI software founder
If you run an AI software company, announcements like Gemini 4 Argon offer two practical takeaways:
First, token pricing continues to bifurcate between generation output and cached input. Building profitable agents requires designing architectures that reuse context aggressively. At $0.10 per million cached tokens, keeping large system guidelines and schemas in memory is economical; generating 1,000,000 output tokens at $10.00 to $20.00 remains an expense you must meter carefully.
Second, upstream model dependencies reinforce why brand, user relationships, and platform distribution matter far more than which model you prompt. A provider can ship a state-of-the-art model on Wednesday and lock it behind an exclusive defense program indefinitely. Founders who build durable products focus on owning their customer base and directory discovery instead of staking their business on immediate access to a single provider's API.
If you are shipping a new software product or agent, you can submit your tool to the SaaSCity startup directory to claim your building on our live city map and earn a backlink. Founders wanting rapid indexing can check out Quick Pass for 24-hour editorial review ($19.99) or Premium for dedicated written launch coverage ($99.99).
Frequently asked questions
Is Gemini 4 Argon released?
Google announced the model on September 30, 2026. However, it is not broadly released to consumers or general API users. It is currently deployed only to vetted cybersecurity defenders in the Fairwind Program.
When will the Gemini 4 Argon API open to developers?
Google has stated the API will open to paid developers and Google AI Ultra subscribers as soon as possible after safety monitoring systems complete verification. Google has not published an exact launch date.
How does Gemini 4 Argon compare to GPT-6 Astra and Claude Opus 5.5?
Argon holds verified leads on DeepSWE v1.1 (77.9%), the Vals Index (68.9%), and AutomationBench (51.3%). It ties GPT-6 Astra on CWE-bench v1 (68%) and Artificial Analysis Intelligence Index (53). However, it trails GPT-6 Astra on FrontierSWE v2 (55.0% vs 65.5%) and Claude Opus 5.5 on Terminal-bench 4.0 (57.4% vs 66.4%).
What is the significance of 1 million output tokens?
Most existing models cap generation output between 32K and 128K tokens. An output window of 1M tokens allows Gemini 4 Argon to conduct hundreds of steps of internal reasoning, perform deep code refactoring, and generate entire multi-file codebases in a single unbroken trajectory.
What is the Fairwind Program?
Fairwind is Google's gated cybersecurity initiative launched on September 2, 2026. It provides verified infrastructure defenders, government teams, and security researchers with access to frontier models without full cyber guardrails so they can find and patch system flaws before attackers exploit them.
Did Google cancel Gemini 3.5 Pro?
Yes. Google originally planned Gemini 3.5 Pro for June 2026, delayed the release internally to improve performance, and skipped the version entirely to launch Gemini 4 Argon as its next flagship.
What are the promotional and standard API prices for Gemini 4 Argon?
Promotional pricing is set at $2.00 per million input tokens, $10.00 per million output tokens, and $0.10 per million cached input tokens (95% discount). Standard pricing after the promotional period will be $4.00 per million input tokens and $20.00 per million output tokens.
Primary sources and reference documentation
- Google DeepMind Official Announcement: Gemini 4 Argon: our next era of frontier intelligence by Koray Kavukcuoglu (September 30, 2026).
- Vals AI Benchmark Portal: Vals Index Economic Impact Leaderboard (September 30, 2026).
- Artificial Analysis: Intelligence Index and Model Economics Tracker (September 30, 2026).
- Google Cybersecurity Initiative: Fairwind Program & Proactive Cyber Defense (September 2, 2026).
- CWE-bench Evaluation: CWE-bench v1 Security Remediation Leaderboard (September 2026).
Get your SaaS in front of founders
List your product on the SaaSCity live city map - a permanent listing, real discovery, and a backlink from a high-DR directory. Free to start; upgrade for a dofollow link and a building on the map.


