谷歌发布了Gemini 4 Argon,这是其最新的尖端模型,旨在缩小与OpenAI和Anthropic等竞争对手的差距,并在关键基准测试中击败其中一些对手。虽然它可能并未在整体表现上遥遥领先,但作为一款尖端模型,至少在其入门价格阶段,它的成本相对较低。
Argon是谷歌在Gemini 3.1 Pro发布七个月以来的首款尖端模型。这使这家广告巨头重新跻身顶级AI实验室前三名,尽管Anthropic可能仍保持领先地位。在经过一段艰难且漫长的开发期后,此前已宣布的Gemini 3.5尖端模型被完全跳过,谷歌如今重返竞争行列。
大多数用户仍需等待
Argon最初将面向“受信任的网络防御者”群体开放,作为Fairwind计划的一部分。这些用户以及谷歌的内部团队将获得没有网络护栏限制的模型版本。谷歌以“分阶段方法”来解释这一逐步推出的策略,认为这一级别的AI能力需要如此。该公司还参与了美国政府的一项自愿计划,该计划允许各机构在模型公开发布前获得访问权限。
早期测试者的反馈将用于完善模型的安全机制。在此之后,谷歌才计划向开发者、企业和消费者开放Argon,首先面向付费API客户和Google AI Ultra订阅用户。公司未给出具体日期,仅表示“尽快”。
定价已确定,至少作为入门价格:每百万输入令牌2美元,每百万输出令牌10美元。缓存输入令牌的成本低95%,相当于每百万约10美分。Gemini 3.8 Flash提供了90%的缓存折扣。这使得谷歌在原始令牌价格方面远低于其他尖端模型,但在令牌消耗量方面并不占优(见下文)。
每百万令牌价格
Gemini 4 Argon(促销价)
Gemini 4 Argon(常规价)
GPT-6 Astra
Claude Fable 5.1
Claude Opus 5.5
输入令牌
2美元
4美元
10美元
10美元
4美元
输出令牌
10美元
20美元
50美元
50美元
20美元
缓存读取操作
0.10美元
0.20美元
1美元
0.25美元
0.20美元
缓存写入操作
不适用
不适用
12.50美元
12.50美元(5分钟)/ 20美元(1小时)
5美元(5分钟)/ 8美元(1小时)
*谷歌未明确说明此数字,但表示缓存比常规输入令牌价格低95%。
谷歌还将输出限制从64,000个令牌提高到一百万个令牌,称其为行业首创。其理念是,如果模型能够在单次推理中生成数十万个令牌,它就能更彻底地思考复杂问题并一次性解决。为了支持这一点,谷歌正在向Gemini API添加一个新的“长解码延续”功能。该功能会暂停长响应,并通过后续请求恢复它们,以防止推理过程因超时而中断。
输入上下文窗口仍保持在一百万个令牌。Argon接受文本、图像、视频和音频作为输入,但仅输出文本。
独立测试显示,Argon的表现与GPT-6 Astra持平,但消耗的令牌更多
Artificial Analysis提供了早期的独立评估。在其可用的最高推理级别“High”下,Gemini 4 Argon在Artificial Analysis智能指数中得分为53分。这使其与OpenAI的GPT-6 Astra(最大值)和Claude Fable 5.1持平,并领先于GPT-6.1 Sol(最大值)一分。
Anthropic的模型仍然领先。Claude Opus 5.5得分为58分,Claude Sonnet 5.5得分为56分。与谷歌上一款尖端模型Gemini 3.1 Pro Preview相比,Argon的得分跃升了23分。“High”通常是Gemini模型的顶级推理层级,尽管有时也支持特殊的“Deep Think”模式。
在当前促销价格下,完成一项智能指数任务的成本为1.99美元。这相当于GPT-6 Astra成本(3.26美元)的60%,但比GPT-6.1 Sol贵2.7倍。一旦折扣结束,成本将上升至3.98美元,比GPT-6 Astra高出约20%。价格优势来自较低的令牌费率,而非效率。Argon平均每个任务消耗62,000个输出令牌,而GPT-6 Astra仅需27,000个。
Artificial Analysis 结果显示,Argon 在“高”档位的表现正迅速逼近第一梯队。| 图片来源:Artificial Analysis
Argon 在代理任务(agentic tasks)方面也取得了显著进步,而根据 Artificial Analysis 的说法,这一直是 Gemini 模型的薄弱环节。在 Artificial Analysis 变体版 AutomationBench-AA 上,它以 77.5% 的得分位居第一,领先于 Claude Sonnet 5.5 (max) 6 个百分点。在 Terminal Bench 4 上,它达到了 57% 的得分,比 Gemini 3.1 Pro Preview 高出 53 个百分点。尽管如此,它仍落后于 Claude Sonnet 5.5(64%)、Claude Opus 5.5(60%)和 GPT-6 Astra(59%)。
Artificial Analysis 还强调了 Argon 较低的幻觉率。在 AA-Omniscience 这一测试事实知识和诚实处理知识盲区的基准测试中,Argon 的幻觉率为 15%。GPT-6 Astra (max) 为 51%,GPT-6.1 Sol (max) 为 54%。Argon 更有可能承认自己不知道答案,而不是错误猜测。然而,其准确率仅为 50%,比 Gemini 3.1 Pro Preview 低 5 个百分点,比 GPT-6 Astra (max, 63%) 低 13 个百分点。在该基准测试的综合得分中,Argon 得分为 42 分,与 GPT-6 Astra(43 分)和 GPT-6.1 Sol(42 分)大致持平。
Google 自身的基准测试结果描绘了一幅更乐观的画面。Argon 在大多数这些基准测试中领先,有时优势巨大。
Google 针对 Gemini 4 Argon 的内部基准测试结果。| 图片来源:Google
Argon 还领跑 Vals Index。根据 Vals AI 的数据,它以 68.9% 的得分位居第一,成为首个登顶该指数的 Gemini 模型。在 22 项测试基准中,Argon 有 20 项进入前五名,尤其在金融、法律、编码和安全领域表现突出。不过,在此测试中,中端模型 Sonnet 5.5 的排名也超过了 Anthropic 的顶级模型 Opus 5.5,因此对此结果需持保留态度。
与往常一样,AI 模型必须在实际使用中证明自身价值,且性能不仅取决于模型本身,还取决于围绕它的软件。这正是 Google 目前的短板:与 Claude Cowork 和 ChatGPT Work 相比,Gemini 应用仍落后于人。
Argon 在人类偏好排名中位居文本类榜首
在 Arena.ai 上,人类会对模型输出进行一对一比较评分,Argon 的表现良好。在 Text Arena 中,Gemini 4 Argon (High) 以 1,525 分位居第一,领先第二名 Claude Opus 4.6 (High) 20 分。这使其成为写作任务的有力竞争者,尤其是在长期缺乏具有竞争力的写作模型、且 Opus 5.5 在人类偏好排名中仍未达到 Opus 4.6 水平的背景下。Google 之前的模型 Gemini 3.8 Flash (High) 曾排在第十一位。
根据 Arena 的数据,Argon 在编码、高难度提示(hard prompts)、指令遵循、长查询和创意写作方面均居首位。它在所有评估的专业领域中也排名第一,并且在英语、中文、俄语以及整体非英语查询中均位列第一。
Web 开发方面的结果则较为平淡。在 Code Arena: WebDev 中,Argon 得分为 1,679 分,排名第八。这比 Gemini 3.8 Flash (High) 提高了 96 分,从第 29 名跃升而来,但未能跻身前列。
在性价比方面,Arena 将 Argon 置于其他模型之上。以每百万 tokens 8 美元的混合费率计算,Argon 推动了 Text Arena 的帕累托前沿(Pareto frontier),目前是排名中成本效益最高的模型。
Google unveiled Gemini 4 Argon, its new frontier model that closes the gap with rivals from OpenAI and Anthropic, beating some of them on key benchmarks. While it may not clearly lead the pack, it is relatively cheap for a frontier model, at least at the introductory price.
Argon is Google's first frontier model in more than seven months, following Gemini 3.1 Pro. It puts the ad giant back among the top three AI labs, though Anthropic likely still holds the lead. After a difficult and drawn-out development period that saw the already-announced Gemini 3.5 frontier model skipped entirely, Google is back in the race.
Most users will have to wait
Argon is initially going to a group of "trusted cyber defenders" as part of the Fairwind program . They and Google's internal teams will get the model without cyber guardrails. Google justifies the gradual rollout with a "phased approach" that AI capabilities at this level require. The company is also taking part in the US government's voluntary program that gives agencies access to new models before public release.
Feedback from early testers will feed into the model's safety mechanisms. Only after that does Google plan to open Argon up to developers, businesses, and consumers, starting with paying API customers and Google AI Ultra subscribers. The company hasn't given a date, saying only "as soon as possible."
Pricing is already set, at least as an introductory rate: $2 per million input tokens and $10 per million output tokens. Cached input tokens cost 95 percent less, working out to about 10 cents per million. Gemini 3.8 Flash had a 90 percent cache discount. That puts Google well below other frontier models on raw token price, though not on token consumption (see below).
Price per million tokens
Gemini 4 Argon (promotional)
Gemini 4 Argon (regular)
GPT-6 Astra
Claude Fable 5.1
Claude Opus 5.5
Input Token
2 $
4 $
10 $
10 $
$4
Dispensing tokens
$10
$20
$50
$50
$20
Cache read operations
$0.10
$0.20
$1
$0.25
$0.20
Cache write operations
N/A
N/A
$12.50
$12.50 (5 min.) / $20 (1 hr.)
$5 (5 min.) / $8 (1 hr.)
*Google doesn't state this figure explicitly but says the cache is 95 percent cheaper than the regular input token price.
Google also raised the output limit from 64,000 to one million tokens, calling it an industry first. The idea is that if the model can generate hundreds of thousands of tokens in a single trajectory, it can think through hard problems more thoroughly and solve them in one pass. To support this, Google is adding a new "Long Decode Continuation" feature to the Gemini API. It pauses long responses and resumes them through follow-up requests so reasoning doesn't hit a timeout.
The input context window stays at one million tokens. Argon accepts text, images, video, and audio as input but only outputs text.
Independent tests show Argon matching GPT-6 Astra but burning more tokens
Artificial Analysis provides an early independent assessment. At its highest available reasoning level, "High," Gemini 4 Argon scores 53 points on the Artificial Analysis Intelligence Index. That ties it with OpenAI's GPT-6 Astra (max) and Claude Fable 5.1, and puts it one point ahead of GPT-6.1 Sol (max).
Anthropic's models still lead. Claude Opus 5.5 sits at 58 points and Claude Sonnet 5.5 at 56. Compared to Google's last frontier model, Gemini 3.1 Pro Preview, Argon jumped 23 points. "High" is typically the top reasoning tier for Gemini models, though a special "Deep Think" mode is sometimes supported as well.
At the current promo price, one Intelligence Index task costs $1.99. That's 60 percent of GPT-6 Astra's cost ($3.26) but 2.7 times more expensive than GPT-6.1 Sol. Once the discount ends, the cost rises to $3.98, about 20 percent above GPT-6 Astra. The price advantage comes from lower token rates, not from efficiency. Argon uses an average of 62,000 output tokens per task, while GPT-6 Astra needs only 27,000.
Artificial Analysis results show Argon at "High" catching up to the top tier. | Image: Artificial Analysis
Argon also made significant gains on agentic tasks, which according to Artificial Analysis have been a weak spot for Gemini models. On AutomationBench-AA, the Artificial Analysis variant, it takes first place at 77.5 percent, six points ahead of Claude Sonnet 5.5 (max). On Terminal Bench 4, it hits 57 percent, a 53-point jump over Gemini 3.1 Pro Preview. That still leaves it behind Claude Sonnet 5.5 (64 percent), Claude Opus 5.5 (60 percent), and GPT-6 Astra (59 percent).
Artificial Analysis also highlights Argon's low hallucination rate. On AA-Omniscience, a benchmark that tests factual knowledge and honest handling of knowledge gaps, Argon's hallucination rate is 15 percent. GPT-6 Astra (max) comes in at 51 percent and GPT-6.1 Sol (max) at 54 percent. Argon is far more likely to admit it doesn't know an answer rather than guess wrong. Its accuracy, however, reaches only 50 percent, five points below Gemini 3.1 Pro Preview and 13 points below GPT-6 Astra (max, 63 percent). On the benchmark's overall score, Argon lands at 42 points, roughly even with GPT-6 Astra (43) and GPT-6.1 Sol (42).
Google's own benchmark results paint a rosier picture. Argon leads in most of those benchmarks, sometimes by wide margins.
Google's internal benchmark results for Gemini 4 Argon. | Image: Google
Argon also leads the Vals Index . At 68.9 percent, it takes first place according to Vals AI, making it the first Gemini model to top the index. Argon finishes in the top five on 20 of 22 tested benchmarks, with particular strength in finance, law, coding, and security. In this test, though, the mid-tier model Sonnet 5.5 also outranks Anthropic's top model Opus 5.5, so take it with a grain of salt.
As always, AI models have to prove themselves in real-world use, and performance depends not just on the model itself but also on the software wrapped around it. That's Google's weak spot right now: compared to Claude Cowork and ChatGPT Work, the Gemini app still lags behind.
Argon tops the human preference rankings for text
On Arena.ai , where humans rate model outputs in head-to-head comparisons, Argon performs well. In the Text Arena, Gemini 4 Argon (High) takes first place with 1,525 points, 20 points ahead of Claude Opus 4.6 (High) in second. That makes it a strong contender for writing tasks, especially after a long drought of competitive writing models and Opus 5.5 still not matching Opus 4.6 in human preference rankings. Google's previous model, Gemini 3.8 Flash (High), had been in eleventh place.
According to Arena, Argon leads in coding, hard prompts, instruction following, longer queries, and creative writing. It also ranks first across all evaluated professional fields, as well as for queries in English, Chinese, Russian, and non-English queries overall.
Web development results are more modest. In Code Arena: WebDev, Argon scores 1,679 points and lands in eighth place. That's a 96-point improvement over Gemini 3.8 Flash (High) and a jump from 29th, but it doesn't crack the top spots.
On price-to-performance, Arena puts Argon ahead of the field. At a blended rate of $8 per million tokens, Argon shifts the Pareto frontier of the Text Arena and is currently the most cost-efficient model in the ranking.
首次收录 · 2026-10-01 · 10.97 分