Anthropic发布了Claude Sonnet 5.5,这是其Claude 5.5系列的第二款模型。它的输出速度提高了30%以上,每项任务的成本降低了多达30%,并且在多个基准测试中几乎与Opus 5.5持平。
Opus 5.5专为需要谨慎判断的复杂任务而设计。Sonnet 5.5则针对定义明确的日常工作,如修复错误、编写文档、构建演示文稿和创建电子表格。Anthropic还宣布将在未来几周内发布Claude Haiku 5.5。该模型将专注于高吞吐量、低成本的用例。
凭借Fable、Opus和Sonnet,Anthropic已经拥有了与OpenAI的GPT-6 Astra、Sol和Luna相对应的产品,后者引发了最新的定价战。粗略来说,Opus略高于Sol,Sonnet高于Luna,而Fable高于Astra,尽管Anthropic的整体收费更高。匹配层级之间的性能差异足够小,成本最终可能成为决定因素。Haiku可以帮助Anthropic缩小这一差距。
编码性能大幅跃升
据Anthropic称,Sonnet 5.5与其前身之间最大的性能差距体现在编码方面。在用于代理式编码测试的Terminal-Bench 4.0上,Sonnet 5.5得分为70.6%,而Sonnet 5仅为10.3%。在重现Cursor编辑器中真实编码会话的CursorBench 4.0上,Sonnet 5.5得分为55.5%,而Sonnet 5为34.1%,仅比Opus 5.5(57.8%)低两分。
Anthropic表示,在FrontierCode 1.1的“High”设置下,Sonnet 5.5的得分比Sonnet 5高出十分之一左右,而每项任务的成本仅为后者的十五分之一。早期测试者称赞该模型理解代码库的速度之快。Sonnet 5.5也比其前身更频繁地批量调用工具,从而减少了所需的步骤数。
推理强度有一个奇怪的细节。在最高努力级别“Max”下,Sonnet 5.5在FrontierCode上的得分实际上低于“Xhigh”级别。Anthropic表示,在最大努力下,该模型更频繁地触发代码审查功能,将工作拆分为多个子代理。在某些情况下,这导致了超时或超出任务范围的变化,而FrontierCode会对这两者进行惩罚。
知识工作几乎与Opus 5.5持平
在GDPval-AA上(这是一个由OpenAI开发的涵盖44个职业和九个行业的知识工作基准测试),Sonnet 5.5得分为1,844分。这几乎与Opus 5.5(1,846分)持平,比Sonnet 5(1,449分)高出约400分。根据Anthropic的数据,OpenAI的GPT-6 Sol得分为1,487分。在视觉图表识别测试Chartography上,Sonnet 5.5的得分从15.6%跃升至61.6%。
| 类别 | 基准测试 | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|---|---|
| 基于代理的编程 | Terminal-Bench 4.0 | 70.6% | 10.3% | 66.4%¹ | — |
| 基于代理的编程 | FrontierCode 1.1 (Main) | 46.2% (Max); 52.1% (Xhigh)² | 42.4% | 54.4% | 49.3% |
| 基于代理的编程 | CursorBench 4.0 | 55.5% | 34.1% | 57.8% | — |
| 知识工作 | GDPval-AA v2.1³ | 1,844 | 1,449 | 1,846 | 1,487 |
| 知识工作 | AA Briefcase v1.1³ | 1,811 | 1,359 | 1,822 | 1,483 |
| 跨学科思维 | Humanity's Last Exam (with tools) | 64.5% | 54.9% | 67.7% | — |
| 计算机技能 | OSWorld 2.1 (partial evaluation) | 80.1% | 57.0% | 81.8% | — |
| 视觉图表识别 | Chartography (without tools) | 61.6% | 15.6% | 64.4% | 53.6% |
¹ 对于Opus 5.5,Terminal-Bench 4.0在“Xhigh”设置下显示最高分。² Sonnet 5.5在FrontierCode上的得分在“Max”级别低于“Xhigh”级别。³ Sonnet 5.5的数值来自一个预发布版本,其中可能影响结构化输出的一个bug后来已被修复。
早期测试者将Sonnet 5.5描述为更自然的对话伙伴,具有设计感。Anthropic声称,该模型能够很好地重写用户界面并执行幻灯片模板,以至于结果几乎不需要微调。
Anthropic 还表示,Sonnet 5.5 是首个仅凭截图就能通关《宝可梦 红》的 Sonnet 系列模型。OpenAI 的 Astra 最近在基准测试中也展示了类似的游戏进展。就在几年前,这类任务仍被视为极具挑战性,通常需要定制算法。如今,同一产品家族中的中端模型也能胜任这些任务。
相同的代币价格,每项任务消耗的代币更少
Sonnet 5.5 每百万代币的价格与 Sonnet 5 相同:输入代币为 2 美元,输出代币为 10 美元,缓存读取为 0.20 美元。Anthropic 表示,由于模型每项任务使用的代币更少,有效成本可降低高达 30%。输出生成速度也提升了 30% 以上。这些说法仍需独立测试来证实。
每百万代币价格
Claude Opus 5.5
GPT-6 Sol
Claude Sonnet 5.5
GPT-6 Luna
输入代币
4 美元
2 美元
2 美元
0.10 美元
输出代币
20 美元
10 美元
10 美元
0.50 美元
缓存读取访问
0.20 美元
0.20 美元
0.20 美元
0.01 美元
缓存写入访问
5 美元
2.50 美元
2.50 美元
0.125 美元
与 Anthropic 的其他模型及 OpenAI 的对应产品一样,Sonnet 5.5 提供可调节的努力程度设置,让用户在成本、速度与输出质量之间进行权衡。Anthropic 表示,在低或中等设置下,Sonnet 5.5 已在多项基准测试中超越 Sonnet 5 的最佳得分,且每项任务的成本仅为后者的约十分之一。然而,为每项任务找到合适的努力程度仍是一门艺术而非科学。
针对更强能力模型的新网络安全保障措施
Claude Sonnet 5.5 现已在所有主要云平台上提供,包括 Amazon Web Services、Google Cloud 和 Microsoft Azure。与 Opus 5.5 和 Sonnet 5 一样,Anthropic 提供零数据保留的模型选项。开发者可以通过 Claude Platform 使用模型 ID “claude-sonnet-5-5” 访问该模型。
由于 Sonnet 5.5 在网络安全方面的能力远超其前身,Anthropic 首次为 Sonnet 系列模型增加了安全保护措施。涉及高风险网络安全任务的请求会被明显重定向至 Sonnet 5。通过扩展的“网络验证计划”,合格的专业人士可以申请分级访问权限。
Anthropic 还针对蒸馏攻击添加了安全分类器,这与它在其最强模型中使用的攻击类型相同。这些措施是否真正有效,将在未来几个月内变得清晰。如果中国实验室曾从这类攻击中受益,关闭这一途径可能会再次拉大与西方实验室之间的差距。
无炒作的人工智能新闻 – 人工策划
订阅 THE DECODER,享受无广告阅读、每周 AI 通讯、每年六次的独家“AI Radar”前沿报告、完整档案访问权限以及评论版块访问权。
立即订阅
Anthropic has released Claude Sonnet 5.5, the second model in its Claude 5.5 family. It generates output more than 30 percent faster, costs up to 30 percent less per task, and nearly matches Opus 5.5 on several benchmarks.
Opus 5.5 is built for complex tasks that demand careful judgment. Sonnet 5.5 targets well-defined everyday work like fixing bugs, writing docs, building presentations, and creating spreadsheets. Anthropic also announced Claude Haiku 5.5 for the coming weeks. That model will focus on high-throughput, low-cost use cases.
With Fable, Opus, and Sonnet, Anthropic already has counterparts to OpenAI's GPT-6 Astra, Sol, and Luna , which kicked off the latest pricing battle. Roughly speaking, Opus sits a bit above Sol, Sonnet above Luna, and Fable above Astra, though Anthropic charges more across the board. Performance differences between matched tiers are small enough that cost may end up being the deciding factor. Haiku could help Anthropic close that gap.
Coding performance jumps sharply
The performance gap between Sonnet 5.5 and its predecessor is most significant in coding, according to Anthropic. On Terminal-Bench 4.0, a test for agentic coding, Sonnet 5.5 hits 70.6 percent compared to Sonnet 5's 10.3 percent. On CursorBench 4.0, which recreates real coding sessions from the Cursor editor, Sonnet 5.5 scores 55.5 percent versus 34.1 percent, landing just two points below Opus 5.5 (57.8 percent).
On FrontierCode 1.1 at the "High" setting, Sonnet 5.5 scores ten points above Sonnet 5 at roughly one-fifteenth the cost per task, Anthropic says. Early testers praised how quickly the model grasps a codebase. Sonnet 5.5 also batches tool calls more often than its predecessor, which cuts the number of steps needed.
Reasoning intensity has one odd wrinkle. At the highest effort level, "Max," Sonnet 5.5 actually scores worse on FrontierCode than at "Xhigh." Anthropic says that at maximum effort, the model more frequently triggers a code-review function that splits work across multiple sub-agents. In some cases, that led to timeouts or changes outside the task scope, both of which FrontierCode penalizes.
Knowledge work nearly matches Opus 5.5
On GDPval-AA, an OpenAI-developed knowledge-work benchmark covering tasks from 44 professions and nine industries, Sonnet 5.5 scores 1,844 points. That nearly matches Opus 5.5 (1,846) and sits about 400 points above Sonnet 5 (1,449). OpenAI's GPT-6 Sol lands at 1,487 by Anthropic's numbers. On Chartography, a visual chart recognition test, Sonnet 5.5 jumps from 15.6 to 61.6 percent.
Category
Benchmark
Sonnet 5.5
Sonnet 5
Opus 5.5
GPT-6 Sol
Agent-Based Programming
Terminal-Bench 4.0
70.6%
10.3%
66.4%¹
—
Agent-Based Programming
FrontierCode 1.1 (Main)
46.2% (Max); 52.1% (Xhigh)²
42.4%
54.4%
49.3%
Agent-Based Programming
CursorBench 4.0
55.5%
34.1%
57.8%
—
Knowledge Work
GDPval-AA v2.1³
1,844
1,449
1,846
1,487
Knowledge Work
AA Briefcase v1.1³
1,811
1,359
1,822
1,483
Interdisciplinary Thinking
Humanity's Last Exam (with tools)
64.5%
54.9%
67.7%
—
Computer Skills
OSWorld 2.1 (partial evaluation)
80.1%
57.0%
81.8%
—
Visual Diagram Recognition
Chartography (without tools)
61.6%
15.6%
64.4%
53.6%
¹ For Opus 5.5, Terminal-Bench 4.0 shows the highest score at the "Xhigh" setting. ² Sonnet 5.5 scores lower on FrontierCode at "Max" than at "Xhigh." ³ Sonnet 5.5 values come from a pre-release version where a since-fixed bug may have affected structured outputs.
Early testers described Sonnet 5.5 as a more natural conversational partner with a feel for design. The model can rework user interfaces and execute slide templates so well that results barely need touch-ups, Anthropic claims.
Anthropic also says Sonnet 5.5 is the first Sonnet model that can play through Pokémon Red using only screenshots. OpenAI's Astra recently showed similar progress on gaming benchmarks . Tasks like these were considered hard just a few years ago and often required custom algorithms. Now even mid-tier models within a product family can handle them.
Same token price, fewer tokens per task
Sonnet 5.5 costs the same per million tokens as Sonnet 5: $2 for input tokens, $10 for output tokens, and $0.20 for cache reads. Because the model uses fewer tokens per task, effective costs drop by up to 30 percent, Anthropic says. Output generation is also more than 30 percent faster. Independent testing still needs to confirm these claims.
Price per million tokens
Claude Opus 5.5
GPT-6 Sol
Claude Sonnet 5.5
GPT-6 Luna
Input tokens
4 $
2 $
$2
$0.10
Output tokens
$20
$10
$10
$0.50
Cache read accesses
$0.20
$0.20
$0.20
$0.01
Cache write accesses
$5
$2.50
$2.50
$0.125
Like Anthropic's other models and OpenAI's counterparts, Sonnet 5.5 offers an adjustable effort setting that lets users trade cost and speed against output quality. At low or medium settings, Sonnet 5.5 already beats Sonnet 5's best scores on several benchmarks at about one-tenth the per-task cost, Anthropic says. Finding the right effort level for each task, though, is still more art than science.
New cybersecurity safeguards for a more capable model
Claude Sonnet 5.5 is available now across all major cloud platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. Like Opus 5.5 and Sonnet 5, Anthropic offers the model with zero data retention. Developers can access it through the Claude Platform using the model ID "claude-sonnet-5-5."
Because Sonnet 5.5 is far more capable in cybersecurity than its predecessor, Anthropic is adding safeguards to a Sonnet model for the first time. Requests involving high-risk cybersecurity tasks get visibly rerouted to Sonnet 5. Through an expanded Cyber Verification Program , qualified professionals can apply for tiered access.
Anthropic has also added safety classifiers against distillation attacks, the same kind used in its most powerful models. Whether these measures actually work should become clear over the coming months. If Chinese labs have been benefiting from such attacks , closing that vector could widen the gap with Western labs again.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Subscribe now
首次收录 · 2026-09-29 · 10.75 分