It’s been a while since we’ve heard about Google’s frontier AI models making headlines. Over the past few weeks, the biggest news came from OpenAI’s GPT-6 Astra, Anthropic’s Claude Opus 5.5, and the new Jev model.
Google’s last model release was four weeks ago with Gemini 3.8 Flash. It was a decent model, but another Flash release wasn’t what most people wanted from the tech giant. With alternatives like DeepSeek V4 Flash and GLM 5.3 already available, another closed model needed to offer a compelling reason to switch.
Today, Google finally addressed the community’s request by releasing a new model on par with Astra/Fable. It’s called Gemini 4 Argon.
Argon focuses on longer tasks that combine reasoning, tool use, and verification.
Here’s a list of all the key changes to Gemini 4 Argon:
- 1M token output limit: The maximum output increases from 64,000 to 1M tokens, giving it substantially more room for reasoning and generation within a single run.
- Stronger software engineering: Argon supports debugging, algorithm design, and large codebase migrations. It reaches 77.9% on DeepSWE v1.1.
- Better execution across business tools: Argon leads AutomationBench at 51.29%, completing workflows that require finding information, updating records, and sending messages across multiple applications.
- Improved visual and long-video understanding: Argon can analyze charts and locate information in long videos. It scores 91.7% on LVBench, ahead of Astra and Opus in Google’s comparison.
- Autonomous vulnerability discovery and patching: Argon can also find software vulnerabilities, validate them, and propose fixes. It scores 68% on CWE-bench v1, tying Astra in Google’s published comparison.
- Stronger defenses and runtime monitoring: Argon adds improved resistance to indirect prompt injection, where malicious instructions arrive through material an agent reads.
Let me repeat the first item: Argon has a 1M output limit. Not a 1M context window.

Argon’s 1M output token limit
That means it can write up to a million tokens in a single response. For context, most frontier models cap output somewhere around 128k.
Google says Gemini 4 Argon agents have already freed over 300 TB of memory across its data centers, with agents already optimizing its own infrastructure.
The agents also analyzed profiling data, found memory optimizations, and applied changes that were subsequently rolled out.
Argon agents are also working on C/C++ to Rust migrations reaching 800,000+ lines of kernel code, with extensive audits and testing before deployment.
These are expensive engineering tasks inside infrastructure Google already operates.
Let’s talk about the benchmarks
Here are three results from Google’s comparison:

Argon’s benchmarks against other frontier models
On Artificial Analysis, Gemini 4 scores 53 on Intelligence Index, matching GPT-6 Astra and trailing Claude Opus 5.5 at 58.

Argon’s benchmarks on Artificial Analysis
These results use Argon’s High reasoning setting, Astra’s Max setting, and Opus’s adaptive Max setting with its default fallback. Higher scores are better.
Here’s more to Artificial Analysis’ findings with Gemini 4 Argon in High reasoning:
- Cost: The 50% launch discount brings its average cost to $1.99 per Intelligence Index task, compared with Astra’s $3.26. Argon actually uses more output tokens (62,000) per task versus 27,000. Its advantage comes from cheaper tokens. At standard pricing, its cost rises to $3.98 per task.
- Agent performance: Argon leads AutomationBench-AA at 77.5% and scores 57% on Terminal-Bench 4. On AA-Briefcase, it achieves the highest recorded rubric pass rate, 65%, but weaker analytical and presentation quality limit its overall result.
- Hallucinations: Its 15% hallucination rate is the lowest among models scoring at least 45 on the Intelligence Index. It is more willing to admit uncertainty, but its 50% accuracy trails Astra’s 63%. Fewer incorrect guesses don’t mean more correct answers.
In terms of Hallucination, Gemini 4 Argon scored 15%, which is insanely low. Grok 4.7 is at 29%, GPT-6 Astra is at 45%, Opus 5.5 at 59%, and Fable 5.1 at 69%.
This is huge! I actually consider hallucination reduction to be one of the biggest upgrades.
The only models below it barely answer anything. None of them get more than 15% right. It gets fewer answers right than Opus 5.5 on max, 50% against 66%. But when it doesn't know, it says so instead of making something up.

Argon’s hallucination benchmarks on Artificial Analysis
Okay, the numbers are impressive on paper. Argon comfortably beats Astra and Opus across these professional-work tests.
But Bloomberg reports that some Google employees find it less effective on actual coding assignments. People with direct access describe uneven performance, and one source identifies front-end design as a weakness.
While Gemini 4 has performed well on benchmarks the industry uses to gauge model efficacy, it does less well when employees actually put it to work, according to people with direct access to the effort. The model struggles to handle certain coding tasks, said the people, who requested anonymity to discuss an internal matter.
You can read the full article here:
Google Grapples Employee Skepticism Gemini 195242680.HtmlBloomberg also reports concerns about benchmaxxing, where development becomes too focused on achieving high test scores. Two people familiar with the model believe Argon is affected by it.
Also, where’s Claude Sonnet 5.5 in the benchmark?
I don’t want model development dictated by whichever benchmark produces the strongest launch chart. I’ve seen it before with their older Pro Gemini models being reported as GOAT in coding, but actually shit in real-world usage.
The professional-work results give Argon a reason to exist. The coding results give me little reason to assume it should replace Astra or Opus.
Should we Trust Google this Time?
Google announced Gemini 3.5 Pro at I/O in May and planned to release it in June. The deadline passed, and the company eventually abandoned the effort.
Sources
- GPT-6 Astragenerativeai.pub
- Claude Opus 5.5generativeai.pub
- Jev modeljimclydemonge.medium.com
- Gemini 4 Argonblog.google
- https://x.com/Google/status/2105388143902175529x.com
- Google’s comparisondeepmind.google
- Artificial Analysisartificialanalysis.ai

