Anthropic正在推出Claude Opus 5.5,这是新系列中的首款模型。该公司表示,该模型在提供与Claude Fable 5.1相当的性能的同时,成本显著更低,运行速度也更快。
据Anthropic称,Opus 5.5在“大多数任务”上与Claude Fable 5.1持平,而运行成本比Opus 5低约40%。Claude Sonnet 5.5和Haiku 5.5预计将在未来几周内推出,同样在性能、效率和安全性方面有所提升。Anthropic的基准测试显示,新款Opus模型在大多数任务上领先于Fable 5.1以及OpenAI成本高昂得多的GPT-6 Astra。
基准/能力
Opus 5.5
Fable 5.1
Opus 5
GPT-6 Astra
GPT-5.6 Sol
代理式编码
Terminal-Bench 4.0[1]
66.4%
55.8%
52.3%
57.9%
37.3%
代理式编码
FrontierCode v1.1 (Main)
54.4%
50.3%
48.0%
53.3%
47.5%
代理式编码
CursorBench 4.0
57.8%
51.8%
46.6%
N/A
41.7%
知识工作
GDPval-AA v2.1
1,846
1,735
1,708
1,542
1,588
业务流程
AutomationBench[1]
40.0%
31.4%
26.9%
41.4%
28.8%
多学科推理
Humanity's Last Exam
67.7%
(带工具)
65.6%
(带工具)
63.6%
(带工具)
57.2%
(带工具)
N/A
代理式科学研究
Terminal-Bench-Science 0.1[1]
58.7%
52.6%
29.0%
64.6%
22.4%
计算机使用
OSWorld 2.0
81.8%
(部分)
80.7%
(部分)
74.0%
(部分)
N/A
N/A
视觉图表识别
Chartography
89.0%
(带工具)
88.4%
(带工具)
83.4%
(带工具)
N/A
N/A
Anthropic表示,新系列主要解决客户在成本、效率和沟通质量方面的反馈,特别是在金融服务、法律和软件开发领域。
降低价格和令牌用量以削减运营成本
Anthropic将Opus 5.5的定价定为每百万输入令牌4美元,每百万输出令牌20美元,而Opus 5的价格分别为5美元和25美元。这意味着令牌价格降低了20%。该公司还将缓存读取成本降低了60%。
Anthropic表示,总运营成本(包括令牌价格和令牌用量)应比Opus 5低约40%。该模型使用的令牌更少,生成输出的速度提高了30%以上。
每百万令牌的单价
Claude Opus 5.5
Claude Opus 5
缓存读取
$0.20
$0.50
输入令牌
$4
$5
输出令牌
$20
$25
缓存写入
$5
$6.25
订阅者的五小时使用限制将增加20%。随着成本的降低,Anthropic表示这些限制总体上延长了25%。用户还可以保留一次限制重置,以备急需之时。
Anthropic正利用编码基准测试来证明其价格与性能的优势。在FrontierCode上,该公司表示Opus 5.5以约20%的任务成本击败了OpenAI的GPT-6 Astra。在Terminal-Bench 4.0上,它声称以40%的成本实现了与Astra相同的性能。在CursorBench上,它表示Opus 5.5以三分之一的成本比GPT-5.6 Sol高出11分。
这次降价是对来自OpenAI以及尤其是中国AI模型压力的回应,后者提供的性能较低,但成本仅为前者的零头。
Anthropic承诺减少“Claude味”
Opus 5.5还被设计为比早期模型沟通更自然。Anthropic表示,它将最重要的信息放在前面,使用较少的行话,并更严格地遵循写作指令。早期测试者形容其写作风格更清晰、更易懂,Anthropic称这使其成为长时间工作会话中更好的合作伙伴。当前的Claude模型因其公式化、晦涩的写作风格而受到大量批评,这种风格有时被称为“Claudish”。
截图来源:PAW
Opus 5.5也是首款在网络安全、生物技术和前沿大语言模型开发方面拥有与Fable 5.1相匹配的安全措施的Opus模型。当这些安全措施触发时,Anthropic表示请求将被透明地路由到其他模型。
用户仍可以在其代码中找到并修复错误,但大多数网络安全任务将交由较旧的Opus 4.8处理。被分类器标记为涉及生物技术或前沿大语言模型开发的请求将交由Opus 5处理。
经核实的组织可以通过生命科学验证计划申请使用该模型进行生物研究。Anthropic 计划在接下来几周内将其现有的网络验证计划扩展至 Opus 5.5。
Claude Opus 5.5 现已在所有平台上提供,包括 Amazon Web Services、Google Cloud 和 Microsoft Azure。使用 Claude Platform 的开发者可以通过模型 ID claude-opus-5-5 访问该模型。
Anthropic 呼吁随着模型能力增强而提高安全标准
Anthropic 计划更严格地筛选强化学习环境。该公司表示,有缺陷的训练环境是导致模型行为偏离对齐的主要来源。该公司还在致力于改进对齐奖励、自动化创建安全训练场景的方法,以及加强安全和监控措施。
AI 实验室正在辩论以多快的速度发布更具能力的模型。OpenAI 最近呼吁为能够自我改进的 AI 系统建立国际标准,而几位研究人员也公开敦促实验室放慢发布速度,直到对齐方法跟上步伐。Anthropic 采取了类似立场,呼吁对可能完全自动化 AI 研究的模型设定更高的安全标准。该公司表示,公共政策应在设定该标准方面发挥更大作用。在 Opus 5.5 发布前,外部组织 Frontier Design 和 METR 对其进行了测试。
新限制措施针对蒸馏攻击并支持欧盟《人工智能法案》合规性
Anthropic 还正在引入针对其所谓“蒸馏攻击”的措施。该公司表示,攻击者利用数千个虚假账户以工业规模提取模型能力,并在没有其安全保护的情况下构建高度强大的模型。Anthropic 引用了一份 2026 年 9 月的威胁报告,记录了其迄今为止已检测并阻止的非法蒸馏活动。
Opus 5.5 发布时带有“保留思维”(Preserved Thinking)功能,这是一种首次随 Fable 5.1 引入的反蒸馏措施。它防止 API 用户编辑 Claude 的先前上下文以提取其推理过程。该措施适用于 2026 年 8 月 31 日或之后创建的 API 账户中的 Fable 5.1 和 Opus 5.5。
Opus 5.5 还包括水印措施,以符合欧盟《人工智能法案》的要求。该模型不再允许在禁用“思维”模式的情况下运行,并且提供零数据保留选项。
Artificial Analysis 将 Opus 5.5 置于其智能指数榜首
独立平台 Artificial Analysis 确认了 Opus 5.5 的强劲表现。在最大努力模式下,该模型在 Artificial Analysis 智能指数中得分 58 分,创下历史新高,比之前的领先者高出数个百分点。它在十项评估中的六项中名列前茅,包括“人类最后一次考试”(Humanity's Last Exam)达到 61.4%(此前最佳:59.1%,Fable 5.1)和 SciCode 达到 66.9%(63.1%,同样为 Fable 5.1)。在 Terminal-Bench 4.0 上,Opus 5.5 得分 59.6%,与 GPT-6 Astra 持平,并比 Opus 5 高出 11 分。
Claude Opus 5.5 以 58 分领先 Artificial Analysis 智能指数,随后是 Claude Fable 5.1 和 GPT-6 Astra,两者均为 53 分。在成本效益图表(底部)中,Opus 5.5 的几个努力水平位于帕累托前沿。| 图片来源:Artificial Analysis
Artificial Analysis 表示,Opus 5.5 使 Anthropic 在 Terminal-Bench 4.0 和 AutomationBench-AA 上与 GPT-6 Astra 持平,同时在代理知识工作方面扩大了领先地位。在私有 AA-Briefcase 基准测试中,Opus 5.5 达到 1,822 的 Elo 分数,比 Fable 5.1 高出 143 分,这是 Anthropic 模型首次在呈现质量上击败 GPT-5.6 Sol。在 GDPval-AA(一个针对现实世界白领工作的 OpenAI 基准测试)中,Opus 5.5 在所有推理模式下均优于 Astra,除了低努力模式。
在GDPval-AA v2.1上,Claude Opus 5.5在几乎所有努力层级下,以相当或更低的任务成本获得了比竞争对手更高的Elo得分。仅在“低”设置下,Astra的表现略胜一筹。| 图片来源:Anthropic
Artificial Analysis确实指出存在效率上的权衡。在最大努力层级下,Opus 5.5每个任务消耗约119,000个输出令牌,远高于Opus 5(73,000)、Fable 5.1(78,000)或GPT-6 Astra(27,000)。较低的令牌价格使其每个任务的成本与Opus 5持平,但并未更便宜。然而,五个Opus 5.5努力层级中有四个位于帕累托前沿,在任务成本指数超过50的范围内,其表现匹配或优于其他所有模型。
没有炒作成分的AI新闻 – 由人工策划
订阅THE DECODER,享受无广告阅读、每周AI通讯、每年六次的独家“AI雷达”前沿报告、完整档案访问权限以及评论板块的参与资格。
立即订阅
Anthropic is launching Claude Opus 5.5, the first model in a new family. The company says it delivers Claude Fable 5.1-level performance while costing significantly less and running faster than its predecessor.
According to Anthropic, Opus 5.5 matches Claude Fable 5.1 "on most tasks" while costing about 40 percent less to run than Opus 5. Claude Sonnet 5.5 and Haiku 5.5 are expected in the coming weeks, with similar gains in performance, efficiency, and safety. Anthropic's benchmarks show the new Opus model ahead of both Fable 5.1 and OpenAI's much more expensive GPT-6 Astra on most tasks.
Benchmark / capability
Opus 5.5
Fable 5.1
Opus 5
GPT-6 Astra
GPT-5.6 Sol
Agentic coding
Terminal-Bench 4.0[1]
66.4%
55.8%
52.3%
57.9%
37.3%
Agentic coding
FrontierCode v1.1 (Main)
54.4%
50.3%
48.0%
53.3%
47.5%
Agentic coding
CursorBench 4.0
57.8%
51.8%
46.6%
N/A
41.7%
Knowledge work
GDPval-AA v2.1
1,846
1,735
1,708
1,542
1,588
Business workflows
AutomationBench[1]
40.0%
31.4%
26.9%
41.4%
28.8%
Multidisciplinary reasoning
Humanity's Last Exam
67.7%
(with tools)
65.6%
(with tools)
63.6%
(with tools)
57.2%
(with tools)
N/A
Agentic scientific research
Terminal-Bench-Science 0.1[1]
58.7%
52.6%
29.0%
64.6%
22.4%
Computer use
OSWorld 2.0
81.8%
(partial)
80.7%
(partial)
74.0%
(partial)
N/A
N/A
Visual chart recognition
Chartography
89.0%
(with tools)
88.4%
(with tools)
83.4%
(with tools)
N/A
N/A
Anthropic says the new series primarily addresses customer feedback on cost, efficiency, and communication quality, particularly in financial services, law, and software development.
Lower prices and token usage cut operating costs
Anthropic prices Opus 5.5 at $4 per million input tokens and $20 per million output tokens, down from $5 and $25, respectively, for Opus 5. That's a 20 percent cut in token prices. The company also reduced cache read costs by 60 percent.
Anthropic says total operating costs, which account for both token prices and token usage , should be about 40 percent lower than Opus 5's. The model uses fewer tokens and generates output more than 30 percent faster.
Prices per 1M tokens
Claude Opus 5.5
Claude Opus 5
Cache reads
$0.20
$0.50
Input tokens
$4
$5
Output tokens
$20
$25
Cache writes
$5
$6.25
Five-hour usage limits for subscribers will increase by 20 percent. With the model's lower costs, Anthropic says those limits stretch 25 percent further overall. Users can also save a limit reset for when they need it most.
Anthropic is using coding benchmarks to make its case on price and performance. On FrontierCode, the company says Opus 5.5 beats OpenAI's GPT-6 Astra at about 20 percent of the cost per task. On Terminal-Bench 4.0, it claims the same performance as Astra at 40 percent of the cost. On CursorBench, it says Opus 5.5 beats GPT-5.6 Sol by 11 points at one-third of the cost.
The price cut is a response to pressure from OpenAI and especially Chinese AI models, which offer lower performance but cost a fraction as much .
Anthropic promises less "Claudish"
Opus 5.5 is also supposed to communicate more naturally than earlier models. Anthropic says it puts the most important information first, uses less jargon, and follows writing instructions more closely. Early testers described its writing as clearer and easier to understand, which Anthropic says makes it a better partner for long work sessions. Current Claude models have drawn plenty of criticism for their formulaic, convoluted writing, sometimes called "Claudish" .
Screenshot via PAW
Opus 5.5 is also the first Opus model with safeguards for cybersecurity, biology, and frontier LLM development that match those of Fable 5.1. When those safeguards kick in, Anthropic says requests are transparently routed to another model.
Users can still find and fix bugs in their code, but most cybersecurity tasks will go to the older Opus 4.8. Requests flagged by classifiers for biology or frontier LLM development will go to Opus 5.
Verified organizations can apply to use the model for biological research through the Life Sciences Verification Program . Anthropic plans to extend its existing Cyber Verification Program to Opus 5.5 in the coming weeks.
Claude Opus 5.5 is available now on all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. Developers using the Claude Platform can access it with the model ID claude-opus-5-5 .
Anthropic calls for higher safety standards as models become more capable
Anthropic plans to screen reinforcement learning environments more strictly. The company says flawed training environments are a major source of misaligned model behavior . It's also working on better alignment rewards, automated ways to create safety training scenarios, and stronger safety and monitoring measures.
AI labs are debating how quickly to release more capable models. OpenAI recently called for international standards for AI systems that could improve themselves , while several researchers have publicly urged labs to slow releases until alignment methods catch up . Anthropic is taking a similar position, calling for a higher safety standard for models that could automate AI research entirely. The company says public policy should have a greater role in setting that standard. External organizations Frontier Design and METR tested Opus 5.5 before its release.
New restrictions target distillation and support EU AI Act compliance
Anthropic is also introducing measures against what it calls distillation attacks . The company says attackers use thousands of fake accounts to extract a model's capabilities at an industrial scale and build highly capable models without its safeguards. Anthropic cites a September 2026 threat report documenting illegal distillation activity it has detected and stopped so far.
Opus 5.5 launches with "Preserved Thinking," an anti-distillation measure first introduced with Fable 5.1. It prevents API users from editing Claude's prior context to extract its reasoning. The measure applies to Fable 5.1 and Opus 5.5 for API accounts created on or after August 31, 2026.
Opus 5.5 also includes watermarking measures to comply with the EU AI Act . The model can no longer run with "Thinking" mode disabled and it is available with Zero Data Retention .
Artificial Analysis puts Opus 5.5 at the top of its intelligence index
Independent platform Artificial Analysis confirms Opus 5.5's strong showing . At max effort, the model scores 58 on the Artificial Analysis Intelligence Index, the highest ever and several points above the previous leader. It leads six of ten evaluations, including Humanity's Last Exam at 61.4 percent (previous best: 59.1 percent, Fable 5.1) and SciCode at 66.9 percent (63.1 percent, also Fable 5.1). On Terminal-Bench 4.0, Opus 5.5 hits 59.6 percent, tying GPT-6 Astra and beating Opus 5 by 11 points.
Claude Opus 5.5 leads the Artificial Analysis Intelligence Index with 58 points, followed by Claude Fable 5.1 and GPT-6 Astra at 53 each. In the cost-performance chart (bottom), several Opus 5.5 effort levels sit on the Pareto frontier. | Image: Artificial Analysis
Artificial Analysis says Opus 5.5 brings Anthropic to parity with GPT-6 Astra on Terminal-Bench 4.0 and AutomationBench-AA while extending its lead in agentic knowledge work. On the private AA-Briefcase benchmark, Opus 5.5 hits an Elo of 1,822, up 143 from Fable 5.1, marking the first time an Anthropic model has beaten GPT-5.6 Sol on presentation quality. On GDPval-AA, an OpenAI benchmark for real-world white-collar work, Opus 5.5 outperforms Astra on every reasoning mode except low.
On GDPval-AA v2.1, Claude Opus 5.5 achieves higher Elo scores than competitors at comparable or lower cost per task across nearly all effort levels. Only at the "low" setting does Astra come out ahead. | Image: Anthropic
Artificial Analysis does flag an efficiency tradeoff. At max effort, Opus 5.5 burns about 119,000 output tokens per task, far more than Opus 5 (73,000), Fable 5.1 (78,000), or GPT-6 Astra (27,000). Lower token prices keep its cost per task in line with Opus 5, but not cheaper. But four of five Opus 5.5 effort levels land on the Pareto frontier, matching or beating every other model above 50 on the index for cost per task.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Subscribe now
首次收录 · 2026-09-23 · 11.21 分