On September 30, 2026, Google officially unveiled Gemini 4 Argon — a new frontier model for complex multi-step tasks: real engineering coding, corporate work in finance and law, and cybersecurity. The main technical feature is the output token limit, raised to 1 million from 64K in the previous generation: an agent can complete a long task in a single trajectory rather than assembling it from short steps. The company claims 77.9% on DeepSWE v1.1 and first place on AutomationBench Zapier and CWE-bench v1, while access is currently limited: expansion will begin with paid API clients and Google AI Ultra subscribers.

image
image

What happened

The announcement, signed by Koray Kavukcuoglu, Senior Vice President of Google DeepMind, was released on September 30, 2026. The company calls Argon a frontier model and for the first time raises the output token limit to 1 million from 64K in the previous generation, allowing the model to output hundreds of thousands of reasoning tokens in a single trajectory. Benchmarks claim 77.9% on DeepSWE v1.1 — SOTA ahead of Claude Opus 5.5 with 74.2% and GPT-6 Astra with 74.1%, first place on AutomationBench Zapier with 51.3%, 91.7% on LVBench for long videos, and a tied first place on CWE-bench v1 for vulnerability patching with a score of 68%. In internal Google projects, Argon agents migrated C/C++ codebases of up to 800K+ lines to Rust, including the Fuchsia Zircon kernel, and also rewrote 32,000 lines of SIMD code in libgav1, achieving a 2.7x speedup over the previous Rust port and freeing up more than 300 TiB of memory. In cybersecurity, the model autonomously finds, verifies, and patches critical vulnerabilities, and a version without cyber-guardrails is available for verified defenders.

Context

The emergence of Argon is a move in the competition among frontier models for long agentic trajectories. In the previous generation, the output limit was 64K tokens, so complex tasks had to be assembled from a sequence of short calls: neither the migration of a large codebase nor an end-to-end vulnerability audit fit in a single pass. Direct competitors in agentic coding are Claude Opus 5.5 from Anthropic and GPT-6 Astra from OpenAI, and Google is openly putting its numbers against them for the first time. A separate layer is trust and safety: Google calls Argon the most resistant model to prompt injection, citing its leadership in the Gray Swan IPI benchmark, and also participates in the voluntary pre-release access process of the US government. As of the announcement, Argon remains a model being tested by a limited circle of users, not a ready-made product for everyone; Google, however, claims to be strengthening protection against abuse.

Why this matters for the industry

For builders of agentic systems, the main shift is not the benchmark table itself, but the unit economics: trajectories of up to 1 million output tokens turn long projects, from code migration and repository analysis to vulnerability audits, from manual labor with hourly pay into a task with predictable token cost, opening up a class of products called "long task as a service." The claimed results on DeepSWE v1.1, AutomationBench, and CWE-bench v1 directly affect the positions of Anthropic and OpenAI in agentic coding and corporate automation, increasing the likelihood of a wave of counter-releases with raised output limits and a new round of re-benchmarking. In cybersecurity, a separate niche is forming: autonomous search, verification, and patching of critical vulnerabilities through the Fairwind program and partnership with Wiz within Scan for Good, where the version without cyber-guardrails is issued only to verified defenders. Teams should already recalculate their unit economics for the final prices of $4/$20 per 1M tokens and design UX for long trajectories: checkpoints, progress display, human-in-the-loop review, and step reproduction.

Why this matters for users

For the reader, Gemini 4 Argon is not yet for everyone: Google is testing the model with a limited circle of users and strengthening protection against abuse, and access expansion will begin with paid API clients and Google AI Ultra subscribers. If you have a budget, it is worth watching for the API opening: the introductory price is $2 per 1M input and $10 per 1M output tokens, cached input is 95% cheaper, and after the introductory period the rate will rise to $4 and $20 per 1M. The practical meaning for a developer is new levels in long tasks: refactoring large codebases, migrations between languages, analysis of long videos, and vulnerability search in a single run instead of a chain of calls. Until access is open, the real benefit is to build your own eval set for your long tasks and calculate the economics in advance for both price scales, so as to compare Argon with your current stack on the first day of the API opening.

What is still unknown / limitations

All benchmark figures (77.9% on DeepSWE v1.1, 51.3% on AutomationBench Zapier, 91.7% on LVBench, 68% on CWE-bench v1) are claimed by Google itself, and the evaluation methodology is not disclosed in available sources: there is no data on pass@k, number of runs, temperature, and contamination control. The gap on DeepSWE v1.1 from Claude Opus 5.5 and GPT-6 Astra is about 3.7 percentage points — on agentic benchmarks, such a difference is easily eaten by the variance between runs. The tied first place on CWE-bench v1 directly indicates the presence of at least one comparable competitor, so the narrative of unconditional leadership is weaker than the headline. Internal cases — the transfer of the Fuchsia Zircon kernel to Rust and the rewriting of SIMD code in libgav1 — rely on internal Google projects and have not been independently verified, and the thesis of hundreds of thousands of reasoning tokens in a single trajectory is currently based on these same cases. Specific deadlines for expanding access beyond the sequence "paid API clients, then Google AI Ultra" are not named in the sources.

Sources

Author

Look at AI, editorial team