Google’s new Gemini 4 Argon is launching to a narrow audience: trusted cyber defenders in Google DeepMind’s Fairwind Program. Developers, enterprises and consumers have no announced access date, even though Google has published API prices and detailed claims about the model’s capabilities.
Google describes Argon as a major advance in long-horizon software engineering, professional work and vulnerability remediation. It also says it needs more time to test safeguards before opening those capabilities to a wider audience. Early third-party benchmarks offer a check on some of its performance claims, but most prospective users cannot yet see how Argon performs in their own systems.
Fairwind Access Comes Before a Public Launch
Google announced Argon on September 30, 2026, and said it was rolling out to a set of trusted cyber defenders through Fairwind. The company is also using the model internally. Neither arrangement amounts to a general Gemini release or an open developer preview.
Google plans to gather feedback from early testers and work on guardrails before expanding access to developers, enterprises and consumers. It has not given a date for that expansion or published a general sign-up path in the announcement. As The Verge’s launch report notes, the immediate release is limited to trusted cyber defenders.
The restriction is tied to what Argon is designed to do. Google says the model can find, validate and patch software vulnerabilities autonomously. It plans to let trusted defenders and its internal teams use those capabilities without cyber guardrails, while continuing to develop protections against misuse for a broader release. That describes a specific exception for vetted defensive work; it does not mean Argon has no safety measures.
Safeguards still under development or testing include defenses against harmful requests and prompt injection, monitoring of model reasoning and actions for misalignment, and more secure environments for high-risk testing. Google also says it is participating in a voluntary U.S. government process for pre-release model access. The announcement presents these steps as reasons for a phased rollout, without giving a timetable for completing one.
Google’s Coding and Enterprise Claims Need Context
Google is positioning Argon as a model that can sustain work across long, multi-step tasks. Its announcement reports a 77.9% score on DeepSWE v1.1, an evaluation of long-horizon software engineering, and a 51.3% score on Zapier’s AutomationBench, which tests execution across business functions. Google calls the latter a first-place result.

The company describes internal work alongside those scores. It says Argon agents are helping migrate C and C++ code to Rust, including work on large codebases. In one example, Google says agents replaced 32,000 lines of SIMD code in an existing Rust port of its libgav1 video decoder and produced a version that runs 2.7 times faster than that Rust port. This is Google’s reported result on a specific project, with no evidence that Argon will deliver comparable speedups on other codebases.
For enterprise work, Google points to financial research, legal tasks and workflow automation. It claims leading results on the Vals Index and domain-specific evaluations for finance and legal work, alongside the AutomationBench score. Those tests address plausible uses for an AI assistant. They do not establish how much supervision a production workflow would need, how often it would fail on unfamiliar material, or whether the economics would work for a particular customer.
A coding agent might produce a credible patch that still fails review; a research agent might complete most steps correctly and mishandle the one source the decision depends on. Such risks grow when a model takes several actions in sequence. Google’s figures are evidence for the tasks and testing conditions it reports, not a measured success rate for everyday deployments.
Cyber Defense Is Both the Selling Point and the Constraint
Cybersecurity gives this release its unusual shape. Google says Argon can discover and remediate vulnerabilities, and reports that it tied for first at 68% on CWE-bench v1, a vulnerability-remediation evaluation. That Google-reported benchmark result is not a general measure of how well the model secures live systems.
Google says Wiz is already using Argon for defensive work through its Scan for Good initiative. It also describes an instance in which the model identified a critical exposure affecting healthcare software. The announcement does not provide enough independent detail to assess that case or predict Argon’s performance across other organizations.
Capable defenders could use stronger tools to find and fix weaknesses sooner. The same ability to investigate software and validate vulnerabilities raises misuse concerns if access is poorly controlled. Google’s approach is to give vetted defenders early access while it tests protections for everyone else. Its effectiveness will depend on both the model’s defensive value and the safeguards Google eventually puts around broader use.
A Million Output Tokens Is Not a Million-Token Context Window
Google says Argon supports an output limit of up to one million tokens, compared with 64,000 tokens for its previous frontier model. Output is what the model generates. A context window describes the material it can take into account. The two limits answer different questions, even when both happen to be described with the same number.
Artificial Analysis lists a separate one-million-token context window for Argon. That should not be mistaken for Google’s output claim. A large context window could let an application supply extensive documents or conversation history; a large output allowance could let an agent continue generating through a long task. Neither specification guarantees that the model will use the available space accurately or economically.
The output figure also has an implementation detail. In its early testing report, Artificial Analysis says it used a Gemini API feature called Long Decode Continuation, which pauses a long response and resumes it across follow-up calls. The advertised limit should not be read as a promise of a single uninterrupted API response containing one million tokens.
API Prices Are Announced, but Most Buyers Cannot Use Them Yet
Google’s introductory API prices are $2 per million input tokens and $10 per million output tokens. It says cached input receives a 95% discount, making the introductory cached-input rate $0.10 per million tokens. After the introductory period, Google plans to charge $4 per million input tokens and $20 per million output tokens. It has not announced when the promotion will end.
These are prices for a planned broader API offering; developers cannot necessarily start sending requests today. They also leave out an important cost variable: how many tokens an agent consumes while completing a task. A lower per-token rate can be offset if a model generates substantially more tokens to get the same work done.
That distinction is particularly relevant to Argon’s large output allowance. Its headline limit describes capacity, not a typical bill or a recommendation to run million-token jobs.
Independent Tests Show Strengths, Not a Settled Ranking
Artificial Analysis provides an early counterpoint to Google’s own performance claims. At its highest reasoning setting, the testing firm reports that Argon scored 53 on its Intelligence Index, matching GPT-6 Astra at its maximum setting and one point ahead of GPT-6.1 Sol at its maximum setting. These are third-party benchmark results under specified settings, not proof that the models are interchangeable in production.
The firm estimates $1.99 per Intelligence Index task for Argon at the introductory prices, compared with $3.26 for GPT-6 Astra in its test. It estimates Argon’s cost would rise to $3.98 per task after the discount ends. Artificial Analysis attributes Argon’s current cost advantage to lower token prices, not frugal generation: it reports an average of roughly 62,000 output tokens per task, versus 27,000 for GPT-6 Astra.
Other results complicate any claim that Argon leads everywhere. Artificial Analysis reports 78% on its AutomationBench-AA evaluation, while Argon’s 57% on Terminal Bench 4 trails several named competitors in that report. AutomationBench-AA’s 78% should not be compared directly with the 51.3% AutomationBench figure in Google’s announcement as if they came from the same test setup.
Artificial Analysis also reports a 15% hallucination rate on AA-Omniscience for Argon, lower than the rates it measured for the leading models it compared. Argon’s 50% accuracy on that evaluation was lower than GPT-6 Astra’s 63%. Together, those figures suggest the model may be more willing to acknowledge uncertainty instead of guessing. They do not show that it answers every question more accurately.
The independent testing lends support to parts of Google’s launch claims, but a large evidence gap remains. With access limited, prospective customers cannot broadly reproduce the findings, measure failure rates on their own workflows, or assess how Google’s eventual public safeguards affect the experience.
Final Thoughts
Google has put forward an unusually specific proposition: Argon can run long, demanding tasks, its introductory token prices are competitive, and its cyber capabilities are valuable enough to release first to vetted defenders. Independent results support parts of that case while showing that price per token, task cost and accuracy can tell different stories.
For now, Argon is a model announcement, not a generally available model release. When Google opens access beyond Fairwind, outside teams will be able to see which capabilities remain available, what safeguards change, and whether they can repeat the reported gains on their own work.
Frequently Asked Questions
4 questions
1Can developers use Gemini 4 Argon now?
No. Google says Gemini 4 Argon is initially rolling out to trusted cyber defenders through its Fairwind Program, while Google also uses the model internally. It plans to expand access to developers, enterprises and consumers after further safeguard testing, but has not announced a public API release date or a general availability date.
2What is Gemini 4 Argon’s one-million-token limit?
Google advertises an output limit of up to one million tokens, meaning the amount the model may generate during extended work. Separately, Artificial Analysis lists a one-million-token context window, which concerns material the model can consider. Artificial Analysis says its long-output testing used a continuation feature that resumes responses across follow-up API calls.
3How much will the Gemini 4 Argon API cost?
Google has announced introductory rates of $2 per million input tokens and $10 per million output tokens, with a 95% discount on cached input. It plans to raise the standard rates to $4 and $20, respectively, after the promotion. Those announced prices do not mean broad developer API access is available yet.
4Is Gemini 4 Argon the best AI model on independent benchmarks?
No single early benchmark establishes that. Artificial Analysis reports that Argon matched GPT-6 Astra at 53 on its Intelligence Index under their respective high-reasoning settings. Its report also shows different relative results on agent tasks and distinguishes a low hallucination rate from answer accuracy. Wider access would allow more teams to test those findings against their own work.
Sources
- Gemini 4 Argonblog.google
- The Verge’s launch reporttheverge.com
- Artificial Analysis listsartificialanalysis.ai
- early testing reportartificialanalysis.ai





