Back to News
ProductAI Understanding briefing

Google unveils Gemini 4 Argon with 1 million token limit and new pricing for trusted cyber defenders

Google announced Gemini 4 Argon, a frontier‑level model with a 1 million‑token output window, introductory pricing of $2 per million input tokens and $10 per million output tokens, and early access limited to a cohort of trusted cyber defenders through its Fairwind program.

5 min readRead the primary source
Source-provided image accompanying Google unveils Gemini 4 Argon with 1 million token limit and new pricing for trusted cyber defenders
Primary-source documentSource recorded
Publisher
blog.google
Source link
blog.googlehttps://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/
Source type
Primary document — an official announcement, paper, filing, or first-party page we read directly.
Also cited

Story last revised

ContextUnderstand this in 60 seconds

Start here

Key terms

Token
A chunk of text processed by language models, such as a word piece or symbol.
API (Application Programming Interface)
A structured way for one software system to send requests to and receive responses from another system.
Prompt Injection
An attack pattern where malicious instructions are inserted into model inputs or retrieved content.
Test yourselfWhat is AI? Quiz

What changed since publication

  1. First published
  2. The blog post adds concrete pricing, expands the token limit to 1 M, provides detailed benchmark scores, and outlines a four‑pillar safety strategy, extending the earlier announcement of Gemini 4 Argon with new operational and safety details.
  3. The VentureBeat report adds detailed benchmark scores, pricing tiers, and rollout strategy to Google’s earlier announcement of Gemini 4 Argon, confirming its 1 million‑token limit and revealing its early‑access program and introductory API rates.
  4. Google’s blog post adds new details to the Gemini 4 Argon launch, including introductory pricing, expanded token limits, benchmark performance numbers, and early cybersecurity use cases, while confirming the model’s limited rollout to trusted cyber defenders.

What happened

Google introduced Gemini 4 Argon, its newest Gemini‑4 family model, on its official blog. The model is being rolled out first to a select group of “trusted cyber defenders” under the Fairwind Program, with plans to expand to developers, enterprises, and consumers later. Argon supports a 1 million‑ output limit—up from the prior 64 K—and is priced at $2 per million input tokens and $10 per million output tokens, with cached input tokens discounted by 95%. The announcement highlights performance gains across software engineering, legal, finance, and cybersecurity tasks, citing benchmark scores such as 77.9% on DeepSWE v1.1, top placement on the Vals Index, 51.3 on AutomationBench, and 91.7 on LVBench. In cybersecurity, Argon can autonomously discover, validate, and patch critical software vulnerabilities, with an early demonstration from Wiz’s Scan for Good program uncovering a high‑risk exposure in healthcare software. Google also details safety measures, including refined guardrails against misuse, prompt‑injection robustness, misalignment monitoring, and hardened sandbox environments.

Google’s blog post announces Gemini 4 Argon as a frontier model designed for deep, long‑horizon reasoning across complex workflows. The model is initially available only to a curated group of trusted cyber defenders via the Fairwind Program, with a roadmap to expand access to developers, enterprises, and consumers.

Argon’s capacity is increased to an industry‑leading 1 million output tokens, a tenfold jump from the previous 64 K limit. This expansion allows the model to generate extensive, coherent outputs in a single pass, supporting tasks like full‑codebase migrations, multi‑step legal drafting, and long‑form video analysis.

Pricing is set at $2 per million input tokens and $10 per million output tokens, with cached input tokens discounted by 95%. These rates are presented as introductory, aimed at encouraging early experimentation while covering compute costs.

Performance benchmarks show Argon leading on DeepSWE v1.1 (77.9% accuracy), the Vals Index for economic impact, AutomationBench (score 51.3), and LVBench (score 91.7). In cybersecurity, Argon ties for first on CWE‑bench v1 with a 68% remediation score, and a demonstration with Wiz’s Scan for Good program uncovered a critical vulnerability in healthcare software.

Google outlines a multi‑layered safety approach: refined misuse detection, robustness against indirect (validated on the Gray Swan IPI benchmark), misalignment monitoring that watches chain‑of‑thought execution, and hardened sandbox environments for training and evaluation.

Source details: blog.google ↗

Why it matters

The 1 million‑ window dramatically expands the horizon for generative AI, enabling single‑prompt reasoning over extremely long contexts such as full codebases, legal contracts, or multi‑hour video analysis. This capability can reduce the need for iterative prompting, lowering latency and cost for complex enterprise workflows. The introductory pricing makes high‑volume token usage more affordable for early adopters, potentially accelerating experimentation in sectors that require large context windows, like finance and cybersecurity. By granting early access to trusted cyber defenders, Google positions Argon as a tool for proactive vulnerability discovery, which could shift how organizations approach threat hunting and patch management. The model’s strong benchmark performance signals a competitive leap over rival offerings, influencing market dynamics and prompting other providers to extend token limits or improve safety mechanisms. However, the limited release and explicit safety framing underscore the ongoing tension between unlocking frontier capabilities and preventing misuse, a balance that will shape regulatory and industry standards.

The 1 million‑ limit removes a major constraint that has limited the applicability of generative models to tasks requiring extensive context, such as analyzing entire code repositories or multi‑hour video footage. This could streamline workflows that currently rely on chunking or multiple prompts, reducing latency and operational overhead.

Introductory pricing lowers the barrier for high‑volume usage, making it financially viable for enterprises to experiment with large‑scale applications that were previously cost‑prohibitive. This may accelerate adoption in sectors like finance, legal services, and cybersecurity where consumption is high.

By targeting trusted cyber defenders first, Google positions Argon as a proactive security tool capable of autonomous vulnerability discovery and patching. If successful, this could shift industry practices from reactive incident response to pre‑emptive defense, potentially reducing the window of exposure for critical infrastructure.

The model’s benchmark dominance signals a competitive shift, pressuring other AI providers to extend limits or improve safety features. This competitive pressure can drive rapid innovation but also raises concerns about an arms race in capability and safety standards.

Google’s detailed safety framework—covering misuse prevention, prompt‑injection robustness, and misalignment monitoring—sets a precedent for responsible rollout of frontier models. The effectiveness of these measures will influence regulatory discussions and industry best practices.

Interactive Mechanism

Interactive Mechanism: How It Actually Works

Explore the underlying technology behind this development interactively.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Interactive Concept Check+10 Points
What is AI? Quiz

Which description best fits "narrow AI", the kind of AI in use today?

What to watch next

Key signals to monitor include the timeline for broader API availability and any adjustments to pricing tiers for enterprise customers. Observers should watch how the Fairwind cohort evaluates Argon’s cybersecurity functions and whether those findings lead to new guardrails or feature expansions. The effectiveness of the model’s safety systems—especially against prompt‑injection attacks and dual‑use threats—will be scrutinized by both the AI research community and policymakers. Additionally, competitor responses, such as extended limits or pricing changes from OpenAI, Anthropic, or other firms, could influence adoption rates. Finally, real‑world case studies from early adopters like Wiz will reveal practical benefits and any unforeseen limitations in large‑scale vulnerability remediation.

When Google expands Argon beyond the Fairwind cohort to broader API customers and Google AI Ultra subscribers, pricing adjustments and tiered access levels will be critical to watch for market impact.

The real‑world performance of Argon in cybersecurity, especially its ability to discover and remediate vulnerabilities without introducing new risks, will be closely examined by security researchers and policymakers.

Competitor responses, such as OpenAI’s or Anthropic’s announcements of larger context windows or new safety mechanisms, could affect Argon’s market positioning and adoption rates.

The robustness of Argon’s safety systems against emerging prompt‑injection techniques and dual‑use scenarios will be a focal point for external red‑team evaluations and academic scrutiny.

Feedback from early adopters, particularly on the model’s cost‑effectiveness and integration challenges in enterprise environments, will shape future product iterations and potential enterprise‑grade offerings.

Related guides & quizzes

What is AI?AI Models ExplainedFuture of AIAI EthicsTest what you know — try a free AI quizLook up an AI term in our glossaryFollow the AI model release tracker

Updates and corrections

This canonical story is updated in place when the developing event materially changes. Its URL and original publication date never change.

  • Google’s blog post adds new details to the Gemini 4 Argon launch, including introductory pricing, expanded token limits, benchmark performance numbers, and early cybersecurity use cases, while confirming the model’s limited rollout to trusted cyber defenders.
  • The VentureBeat report adds detailed benchmark scores, pricing tiers, and rollout strategy to Google’s earlier announcement of Gemini 4 Argon, confirming its 1 million‑token limit and revealing its early‑access program and introductory API rates.
  • The blog post adds concrete pricing, expands the token limit to 1 M, provides detailed benchmark scores, and outlines a four‑pillar safety strategy, extending the earlier announcement of Gemini 4 Argon with new operational and safety details.
See the public corrections log
Found this useful?