首页
创新与人工智能
模型与研究
Gemini 模型
Gemini 4 Argon:开启前沿智能的新纪元
2026年9月30日
阅读时间约9分钟
Gemini 4 Argon 在现实世界的软件工程、法律和金融等企业知识工作以及网络安全防御等复杂工作流程中,提供了前沿的性能表现。
Koray Kavukcuoglu
Google DeepMind 高级副总裁兼 Google 首席人工智能架构师
分享
收听文章
10分35秒
阅读 AI 生成的摘要
本文目录
介绍 Gemini 4 Argon
改变我们在 Google 的工作和构建方式
在解决最复杂的问题上加倍努力
赋能跨领域的编码和企业工作流程
在防御性网络安全方面处于领先地位
在广泛可用之前加强前沿安全保障
即将推出
今天,我们宣布推出全新的前沿模型 Gemini 4 Argon,该模型正通过我们的 Fairwind 计划向一批值得信赖的网络安全防御者开放。Argon 旨在在复杂、长周期的工作流程中保持深度推理能力,从根本上改变了我们在 Google 的工作和构建方式。它在现实世界的软件工程、法律和金融等企业知识工作以及网络安全防御等复杂工作流程中提供了前沿的性能表现。
以这种水平安全地发布前沿功能需要采取分阶段的方法。我们正积极参与美国政府自愿进行的模型发布前访问流程,同时逐步扩大访问范围。在迭代护栏(guardrails)之前,我们将继续收集早期测试者的反馈,以便尽快向开发者、企业和消费者提供 Argon。
Argon 将以入门级价格推出
1
输入令牌每百万个2美元,输出令牌每百万个10美元,缓存的输入令牌价格为输入令牌价格的5%。
改变我们在 Google 的工作和构建方式
Gemini 4 Argon 已经在我们的内部工作流程中发挥作用,数千名 Google 员工强调了该模型在专业编码任务、进行更深入的研究以及写作质量方面的优势。它正在帮助团队更快地构建产品,推动工程生产力的边界,并加速突破:
量子算法优化:Argon 正在帮助我们的量子计算研究人员优化对重要应用程序造成瓶颈的子例程的时空资源(量子比特数 × 门数)。在一个例子中,它在几分钟内就将表现超越了已发布的基线40%。
内存效率:一组 Argon 智能体分析了全公司的性能分析遥测数据,自主识别并应用了 Google 数据中心的内存优化措施。一旦部署,将释放超过300 TiB的内存,预计总节省量在500 TiB到1 PiB之间。
大规模代码库迁移与优化:Argon 智能体正在 Google 范围内将 C/C++ 代码库迁移至 Rust,规模从 re2、libgav1 等核心库中的数万行代码扩展到 Fuchsia Zircon 内核的80多万行代码。鉴于许多这些系统的关键性,此类大规模重写在部署到生产环境之前,都要经过严格的自动化和人工审计、仿真测试以及审查。
对于 Google 用于视频解码的开源软件 libgav1,Argon 智能体利用现有的 Rust 端口,通过运行多轮基于性能分析的实验,研究编译器的输出,从而生成了安全的 Rust 代码,使编译器能够自动对其进行向量化处理。最终结果是一个内存安全的视频解码器,其运行速度比 Rust 端口快2.7倍,且视频输出完全相同,使其更接近于优化后的 C++ 版本。
在解决最复杂的问题上加倍努力
为了支持 Gemini 4 Argon 在更长、更复杂用例中的能力,我们大幅扩展了模型的输出令牌限制,达到行业领先的 100 万(1M)个令牌,此前为 6.4 万个(64K)令牌。当模型有足够的空间进行深入思考并在单次轨迹中生成数十万个令牌时,它会在推理过程中增加一个新的深度层次,从而一次性解决棘手的问题。
赋能跨领域的编码和企业工作流
Gemini 4 Argon 在编码、推理和多模态方面的能力,以及其维持长周期、多步骤任务的能力,使其能够在一系列企业工作流中脱颖而出。
Google 工程师已将 Argon 用于他们的日常任务,从日常调试到大规模代码库迁移和算法设计。它在 DeepSWE v1.1(77.9%)上树立了新的最先进水平(state of the art),该基准测试衡量模型在现实世界长周期软件工程任务中的表现。
除了编码之外,Argon 还是 Vals Index 上的领先模型,该指数衡量金融、编码、法律和税务工作 across 的经济影响,每个领域都根据其对美国 GDP 的贡献进行加权。在其他特定领域的评估中,我们也看到了同样领先的性能,例如 Vals Finance Agent v2(多步金融研究)和 Harvey 的法律代理基准测试(法律研究和起草)。在 AutomationBench(Zapier 衡量核心业务功能端到端执行的基准测试)上,Argon 以 51.3% 的得分排名第一。
当知识工作需要视觉理解时,Argon 也表现出独特的优势。它能够驱动专业的图表分析,从长视频中识别细节,并根据一系列文档采取行动。例如,在衡量长视频理解的 LVBench 上,Argon 以 91.7% 的得分处于最先进水平。
防御性网络安全领域的领先者
为了更好地装备网络防御者应对新一代网络攻击,我们训练 Gemini 4 Argon 具备强大的网络安全防御能力。Argon 能够自主发现、验证并修补关键软件漏洞。对于受信任的防御者以及 Google 内部团队,我们将发布不带网络安全护栏(guardrails)版本的 Argon,以便他们充分利用其前沿级别的网络安全防御能力。
Wiz 已通过其“Scan for Good”计划使用 Argon 进行网络安全防御——该项目致力于通过发现并修复高风险暴露来免费保护关键公共基础设施。在早期展示其影响力的演示中,该模型发现了一个关键漏洞,该漏洞导致全球医院使用的医疗软件中的敏感个人信息面临风险,识别出此前其他前沿模型所忽略的严重风险。
在评估模型修复安全漏洞能力的 CWE-bench v1 上,Argon 以 68% 的最高得分并列第一,这是在 CWE-bench v0 上 3.8 Flash Cyber 的前沿性能基础上的进一步提升。
Gemini 4 Argon 在漏洞发现方面相比 3.8 Flash Cyber 取得了令人印象深刻的飞跃。例如:
在 Google 的内部综合漏洞基准测试中,Argon 发现了跨越 20 种编程语言的复杂代码库中的广泛暴露点。
在 Wiz 的内部黑盒渗透测试基准测试中(该测试评估模型在没有源代码的情况下分析实时 Web 系统的能力),Argon 在发现攻击面、识别漏洞以及生成用于验证的概念证明证据方面优于 3.8 Flash Cyber。
在广泛可用之前加强前沿安全措施
在广泛推出 Gemini 4 Argon 之前,我们继续加强四个主要领域的关键前沿安全措施:
防御滥用:为防止不良行为者利用 Argon 发动网络或化学、生物、放射性和核(CBRN)攻击,根据我们的前沿安全框架,该模型被设计为拒绝有害请求,同时保留合法的、军民两用的科学研究。我们正加强此次发布中安全措施鲁棒性,包括改进监控模型内部激活状态的技术,以发现滥用行为。这些安全措施已通过内部和外部红队使用手动和自动攻击方法的组合进行了鲁棒性测试。
防御提示注入攻击:Argon 也是我们对抗间接提示注入最具韧性的模型,此类攻击利用恶意指令或上下文来劫持模型行为。这些是复杂的攻击,需要持续的警惕和多层次的防御。通过自动化红队和对抗性训练,Gemini 4 Argon 在 Gray Swan 的间接提示注入(IPI)基准测试中处于提示注入鲁棒性的领先地位。
监控对齐偏差:为了防止 Argon 越界,以超出用户意图的方式尝试完成任务,我们正在部署对齐偏差缓解措施,监控 Argon 的思维链和行动,并在必要时停止执行。
我们使用了类似的系统来监控我们的训练运行,并向专门的应急响应团队发送警报,采取谨慎预防措施,避免将调查结果反馈回训练过程,以免冒险让 Argon 的推理能力学会规避我们的监控。我们强烈鼓励行业其他同仁在这些关键能力提升期保持推理透明度,以应对对齐风险,从而使模型思维在识别和诊断对齐偏差方面继续发挥积极作用。
加固系统:随着前沿模型能力日益增强,对其安全测试需要能够跟上系统本身发展的安全环境。按照我们的智能体控制路线图,我们正在加固沙盒环境,在高危训练或评估开始前对其进行隔离和密封。我们致力于与合作伙伴分享这些智能体安全最佳实践,以提升整个生态系统的安全性。
即将推出
我们构建了具备前沿级能力的 Gemini 4 Argon,涵盖编码、知识工作、网络安全防御和创意写作,旨在成为开发者、专业人士和企业应对最棘手问题的合作伙伴。我们感谢首批网络防御者和受信任测试人员,他们的真实世界评估和反馈将帮助我们在向开发者、企业和消费者发布前(首先面向付费 API 客户和 Google AI Ultra 订阅用户)进一步强化我们的系统。
发布于:
Gemini 模型
Home
Innovation & AI
Models & research
Gemini Models
Gemini 4 Argon: our next era of frontier intelligence
Sep 30, 2026
9 min read
Gemini 4 Argon delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense.
Koray Kavukcuoglu
SVP, Google DeepMind and Chief AI Architect, Google
Share
Listen to article
10:35 minutes
Read AI-generated summary
In this article
Introducing Gemini 4 Argon
Changing how we work and build at Google
Working harder on your most complex problems
Enabling coding and enterprise workflows across domains
Leading in defensive cybersecurity
Strengthening frontier safeguards before broad availability
Rolling out soon
Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program. Built to sustain deep reasoning across complex, long-horizon workflows, Argon is fundamentally changing the way we work and build at Google. It delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense.
Safely releasing frontier capabilities at this level requires a phased approach. We are actively engaged in the U.S. government’s voluntary process for pre-release model access while we gradually expand access. We’ll continue to gather feedback from early testers as we iterate on guardrails before making Argon available to developers, enterprises, and consumers as soon as possible.
Argon will launch at an introductory price
1
of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price.
Changing how we work and build at Google
Gemini 4 Argon is already powering our internal workflows, with thousands of Googlers highlighting the model’s strengths in specialized coding tasks, conducting deeper research, and writing quality. It’s helping teams build faster and push the boundaries of engineering productivity and accelerating breakthroughs:
Quantum algorithmic optimization: Argon is helping our quantum computing researchers optimize the spacetime resources (qubits × gates) of subroutines that bottleneck important applications. In one example, it beat the published baseline by 40% in a matter of minutes.
Memory efficiency: A team of Argon agents analyzed fleet-wide profiling telemetry to autonomously identify and apply memory optimizations across Google’s data centers, freeing up over 300 TiB of memory once rolled out, with an estimated 500 TiB to 1 PiB in total savings.
Large Scale Codebase Migrations and Optimizations: Argon agents are working on migrating C/C++ codebases to Rust across Google — scaling from tens of thousands of lines in core libraries like re2, libgav1 up to 800K+ lines for the Fuchsia Zircon kernel. Given the criticality of many of these systems, such large-scale rewrites are undergoing rigorous automated and manual auditing, emulation testing, and review before rolling out to production.
For libgav1, Google's open source software for decoding video, Argon agents took an existing Rust port and replaced 32K lines of SIMD code by running many rounds of profile-guided experiments, studying the compiler's output, producing safe Rust so the compiler would vectorize it automatically. The end result is a memory-safe video decoder that runs 2.7x faster than the Rust port, with identical video output, bringing it closer to the optimized C++.
Working harder on your most complex problems
To support Gemini 4 Argon’s capabilities across longer, more complex use cases, we are significantly expanding the model’s output token limit to an industry-leading 1M tokens, up from the previous 64K tokens. When the model has the headroom to think deeply and generate hundreds of thousands of tokens in a single trajectory, it adds a new level of depth in reasoning to solve tough problems in one go.
Enabling coding and enterprise workflows across domains
Gemini 4 Argon’s capabilities across coding, reasoning, and multimodality and its ability to sustain long, multi-step tasks enable it to excel across a range of enterprise workflows.
Google engineers have been using Argon for their daily tasks, from everyday debugging to large-scale codebase migrations and algorithm designs. It sets a new state of the art on DeepSWE v1.1 (77.9%), which measures a model’s performance in real-world long-horizon software engineering tasks.
Beyond coding, Argon is the leading model on the Vals Index, which measures economic impact across finance, coding, legal, and tax work, with every sector weighted by its contribution to U.S. GDP. We see similarly leading performance across other domain specific evaluations, like Vals Finance Agent v2 (multi-step financial research) and Harvey’s Legal Agent Benchmark (legal research and drafting). On AutomationBench, Zapier’s benchmark measuring end-to-end execution across core business functions, Argon ranks #1 with a score of 51.3%.
Argon is also uniquely strong when knowledge work requires visual understanding. It’s able to drive professional chart analysis, identify details from long videos, and take action based on a series of documents. For example, on LVBench, which measures long video understanding, Argon is state of the art with a score of 91.7%.
Leading in defensive cybersecurity
To better equip cyber defenders for the new era of cyberattacks, we trained Gemini 4 Argon to be highly capable at cybersecurity defense. Argon can autonomously find, validate, and patch critical software vulnerabilities. For trusted defenders and our own internal teams at Google, we’ll be releasing Argon without cyber guardrails so they can leverage its full frontier-level cybersecurity defense capabilities.
Wiz is already using Argon for cybersecurity defense through its Scan for Good initiative – a program dedicated to protecting critical public infrastructure for free by finding and remediating high-risk exposures. In an early demonstration of its impact, the model uncovered a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide, identifying a severe risk that previous frontier models had missed.
On CWE-bench v1, which evaluates the model’s ability to remediate security vulnerabilities, Argon ties for first place with a top score of 68%, building on 3.8 Flash Cyber’s frontier performance on CWE-bench v0.
Gemini 4 Argon demonstrates impressive leaps in vulnerability discovery over 3.8 Flash Cyber. For example:
On Google’s internal comprehensive vulnerability benchmark, Argon uncovered a wide range of exposures across complex codebases spanning 20 programming languages.
On Wiz’s internal black-box penetration testing benchmark, which tests a model’s ability to analyze live web systems without source code, Argon outperforms 3.8 Flash Cyber in discovering the attack surface, identifying vulnerabilities, and producing proof-of-concept evidence to validate them.
Strengthening frontier safeguards before broad availability
Before rolling out Gemini 4 Argon broadly, we’re continuing to strengthen critical frontier safeguards across four main areas:
Defending against misuse: To prevent bad actors from using Argon for cyber or chemical, biological, radiological, and nuclear (CBRN) attacks, the model is designed to refuse harmful requests while preserving legitimate, dual-use scientific research, as per our Frontier Safety Framework. We are strengthening the robustness of our safeguards for this launch, including improving our techniques to monitor the model’s internal activations to spot misuse. These safeguards underwent robustness testing by internal and external red teams using a combination of manual and automated attack methods.
Defending against prompt injection attacks: Argon is also our most resilient model yet against indirect prompt injections, where malicious instructions or context are used to hijack a model’s behavior. These are complex attacks that require constant vigilance and multiple layers of defense. Through automated red teaming and adversarial training, Gemini 4 Argon is leading in prompt injection robustness on the Gray Swan’s Indirect Prompt Injection (IPI) benchmark.
Monitoring for misalignment: In order to prevent Argon from stepping out of bounds to try to accomplish a task in a way that goes beyond the user’s intentions, we are deploying misalignment mitigations that monitor Argon’s chain-of-thought and actions and stop execution when necessary.
We used a similar system to monitor our training runs and send alerts to a dedicated incident response team, taking careful precautions against feeding the findings back into training so as to not risk shaping Argon’s reasoning to evade our monitoring. We strongly encourage the rest of the industry to preserve reasoning transparency in these pivotal moments of increased capabilities while navigating alignment risks, so that model thoughts remain helpful in identifying and diagnosing misalignment.
Hardening systems: As frontier models grow increasingly capable, safely testing them requires secure environments that can keep up with the systems themselves. In line with our agent control roadmap, we are hardening our sandboxed environments by isolating and sealing them before high-risk training or evaluations begin. We’re committed to sharing these agent security best practices with our partners to improve security across the ecosystem.
Rolling out soon
We built Gemini 4 Argon with frontier-level capabilities on coding, knowledge work, cybersecurity defense, and creative writing to be a partner for developers, professionals, and enterprises while they tackle the most difficult problems. We’re grateful for the initial cohort of cyber defenders and trusted testers whose real-world evaluations and feedback will help us strengthen our systems before we release to developers, enterprises, and consumers, starting with paid API customers and Google AI Ultra subscribers.
Posted in:
Gemini models
首次收录 · 2026-10-01 · 11.86 分