Google Introduces Gemini 4 Argon for Advanced Coding, Enterprise Work and Cybersecurity

Google has announced Gemini 4 Argon, its latest and most advanced artificial intelligence model, as the company pushes deeper into complex software development, enterprise work and cybersecurity.

The new model is the first release in Google’s Gemini 4 series. Google says Gemini 4 Argon is built to handle long, multi-step tasks that require sustained reasoning rather than producing a quick answer to a single prompt.

The company is initially making Argon available to a limited group of trusted cybersecurity professionals through its Fairwind Program. A wider release for developers, businesses and consumers will come later, with Google saying it wants to gather more feedback and strengthen its safety systems before expanding access.

The launch comes as Google competes with OpenAI and Anthropic for the next generation of advanced AI models. Reuters reported that Google’s latest release follows delays in its AI model roadmap and comes as the company seeks to compete more directly with rival frontier models.

Gemini 4 Argon Can Handle Longer and More Complex AI Tasks

One of the biggest changes in Gemini 4 Argon is its much larger output limit. Google has increased the model’s output capacity from 64,000 tokens to 1 million tokens. This allows Argon to continue working through very long tasks in a single trajectory instead of stopping after a relatively short response and requiring the user or another system to restart the process.

The change is particularly relevant for tasks such as software engineering, research, legal analysis and financial work, where an AI system may need to process information, reason through several steps and produce a large amount of output.

Google says Argon is intended to maintain its reasoning across these longer workflows. The company has already been using the model internally for coding, debugging, research and large-scale engineering projects.

That focus marks a shift from AI models being used mainly for individual questions or short pieces of content toward systems that can work through larger projects over an extended period.

ALSO READ: Anthropic Launches Claude Opus 5.5 With Fable-Level Performance at 60% Lower Cost

Google Is Already Using Gemini 4 Argon on Major Engineering Projects

Google Is Already Using Gemini 4 Argon on Major Engineering Projects

Google is testing Gemini 4 Argon across several areas of its own operations. According to the company, thousands of Google employees are already using the model for specialized coding tasks, research and writing. Google also highlighted several internal projects where Argon agents have been used to work on engineering problems.

In one example, Google said Argon helped its quantum computing researchers optimize a computational bottleneck and beat a published baseline by 40% in a matter of minutes. The company also used Argon agents to analyze data-center performance information and identify memory optimizations. 

Google estimates that these changes could free more than 300 TiB of memory once deployed, with total potential savings estimated at between 500 TiB and 1 PiB. Argon is also being used for large codebase migrations. Google said its agents are helping move C and C++ code to Rust, including projects ranging from tens of thousands of lines of code to more than 800,000 lines in the Fuchsia Zircon kernel.

Google said these large migrations are being subjected to automated and manual audits, testing and reviews before they reach production. That is significant because the model is being used on software that forms part of Google’s broader infrastructure rather than only on experimental coding projects.

Gemini 4 Argon Shows Strong Performance on Coding and Enterprise Tests

Google is also using software engineering benchmarks to demonstrate Argon’s capabilities. The model scored 77.9% on DeepSWE v1.1, a benchmark focused on real-world, long-horizon software engineering tasks. Google described the result as a new state-of-the-art score on the benchmark.

Argon also performed strongly on enterprise-focused evaluations. Google said it leads the Vals Index, which measures performance across finance, coding, legal and tax-related work.

On AutomationBench, a benchmark from Zapier that measures end-to-end execution across business functions, Argon scored 51.3%. Google also reported a 91.7% score on LVBench, which measures long-video understanding.

These figures come from Google’s own evaluation and should therefore be viewed in that context. Other benchmark comparisons show that rival models still perform better on some individual coding and terminal-based tasks.

Google Puts Cybersecurity at the Center of Gemini 4 Argon

Google is putting particular attention on Argon’s cybersecurity capabilities. The company says the model can autonomously find, validate and patch critical software vulnerabilities. For its trusted cybersecurity partners and internal security teams, Google plans to provide access without some of the cyber safeguards applied to broader releases so those teams can use the model’s full defensive capabilities.

Google said cybersecurity company Wiz is already using Argon through its Scan for Good initiative, which focuses on identifying and fixing security problems affecting critical public infrastructure.

In one early test, Google said Argon discovered a critical vulnerability in healthcare software that could expose sensitive personal information. The company said previous frontier models had failed to identify the vulnerability.

Argon also recorded a 68% score on CWE-bench v1, a benchmark that measures an AI model’s ability to remediate software vulnerabilities. Google said the result ties for the top score on that benchmark.

Google Tests Gemini 4 Argon Against Cyber and Other AI Risks

Google Tests Gemini 4 Argon Against Cyber and Other AI Risks

Google says it is continuing to strengthen Gemini 4 Argon’s safety systems before making the model broadly available. The company highlighted several areas, including protection against cyber misuse and chemical, biological, radiological and nuclear-related misuse. Google said it has also tested the model through internal and external red-team exercises.

Prompt injection is another focus. These attacks attempt to manipulate an AI system by placing malicious instructions in content the model processes. Google said Argon has been trained and tested to improve its resistance to indirect prompt injection attacks. 

The company also says it is using systems to monitor the model’s actions and reasoning for signs that it could move beyond the user’s intended objective. Google is also hardening the environments in which its frontier models are tested. The company said secure, isolated environments are becoming increasingly important as AI systems gain the ability to perform more complex tasks.

ALSO READ: Anthropic and OpenAI Push for Stronger Rules as AI Safety Concerns Grow

Gemini 4 Argon API Pricing Starts at $2 per Million Input Tokens

Google has also announced initial API pricing for Gemini 4 Argon. The introductory price will be $2 per million input tokens and $10 per million output tokens. Cached input tokens will receive a 95% discount from the standard input-token price.

After the introductory period, Google says the price will rise to $4 per million input tokens and $20 per million output tokens. The pricing is aimed at developers and businesses that want to integrate the model into their own applications and workflows rather than only use it through a consumer chatbot.

Gemini 4 Argon Is Not Yet Available To Everyone

Despite the launch announcement, most users cannot access Gemini 4 Argon yet. Google is first rolling out the model to trusted cyber defenders through its Fairwind Program. The company is also participating in the U.S. government’s voluntary process for pre-release access to advanced AI models.

Google says the next stage will bring Argon to paid API customers and Google AI Ultra subscribers, followed by broader access for developers, enterprises and consumers. The company has not provided a specific date for full public availability.

This phased approach reflects the growing concern around AI models that can independently perform tasks such as writing and modifying software, finding vulnerabilities and carrying out long sequences of actions.

Gemini 4 Argon Raises the Stakes in Google’s AI Competition

Gemini 4 Argon Raises the Stakes in Google’s AI Competition

The Gemini 4 Argon launch gives Google a new flagship model as competition among leading AI companies continues to intensify. OpenAI, Anthropic and Google are increasingly focused on models that can do more than answer questions. 

Their latest systems are being developed to handle software projects, research, business operations and other tasks that can require multiple steps and extended reasoning. Argon’s 1-million-token output limit, coding capabilities and focus on cybersecurity are central to Google’s pitch for the new model. The company is keeping the initial release limited while it tests the system with trusted users and continues work on its safety controls.

For now, Gemini 4 Argon remains a limited-access model. Its broader release will give developers and businesses a better opportunity to test whether Google’s claims about long-running AI workflows translate into practical improvements outside the company’s own testing environment.