我们推出 Claude Opus 5.5,这是我们要推出的全新 Claude 5.5 系列中的第一款模型。它在大多数任务上的表现达到了 Claude Fable 5.1 的水平,而运行成本比 Opus 5 低 40%。
Claude Opus 5.5 是我们呼吁“放缓前沿发展”以来的首款发布产品。在发布前,它已由包括 Frontier Design 和 METR 在内的外部评估人员进行了测试。在我们执行的自动化行为审计——这是我们所做的最全面的对齐测试中——Opus 5.5 是我们迄今测试过的表现最强的模型。它还配备了我们为最具能力模型开发的安全保障措施。
以下是您从 Opus 5.5 中可以期待的一些改进:
性能。Opus 5.5 相比 Opus 5 是一个巨大的飞跃。它是新的领先模型,早期测试者在处理最复杂的工作时看到了性能的显著提升。一位测试者在不到一天的时间内完成了一项涉及 68 万行代码的迁移工作——这项工作原本需要工程团队花费数周时间。它擅长发现并修复软件中的低效问题:当我们要求它缩短 Web 应用每个页面的加载时间时,Opus 5.5 在 40 次尝试中成功了 39 次,而 Opus 5 虽然也做出了一些改进,但幅度较小且改变了应用的行为。另一位测试者让多个 Claude 模型根据单个提示构建游戏;在图形的强度和打磨程度上,Opus 5.5 的得分高于任何其他模型。
安全。Opus 5.5 在我们自动化行为审计中取得了迄今为止任何模型的最佳分数,该审计是我们用于在数千个模拟场景中测试 Claude 的对齐套件。与最近的模型相比,它采取难以逆转的行动或超出给定边界行事的可能性要小得多,并且比 Opus 5 更能抵抗提示注入。我们还扩大了对齐测试的范围,以涵盖更长的任务、不可能的任务以及基于真实事件构建的场景,尽管它仍有局限性。我们评估的详细信息可在 Opus 5.5 System Card 中找到。
由于 Opus 5.5 在生物学和网络安全方面的能力与 Claude Mythos 5.1 相当,因此我们将其部署时采用了与 Claude Fable 5.1 类似的安全保障措施。经过验证的组织今天可以申请我们的生命科学验证计划(Life Sciences Verification Program),以使用 Opus 5.5 进行生物学研究。在未来几周内,我们还将扩大对网络安全验证计划(Cyber Verification Program)的访问权限,经认证的网络安全从业者将能够将其用于工作。
成本与速度。Opus 5.5 提供服务所需的计算资源少于 Opus 5,其定价也反映了这一点。我们的测试显示,在默认设置下,它在典型负载下的成本比 Opus 5 低 40%。输入和输出令牌的价格分别为每百万个 4 美元和 20 美元,比 Opus 5 低 20%。缓存读取(占代理和编码工作成本的大部分)为每百万个令牌 0.20 美元,比 Opus 5 低 60%。Opus 5.5 的输出生成速度也比 Opus 5 快 30% 以上。
除了降价之外,我们还提高了 Pro、Max、Team 以及按席位计费的 Enterprise 计划的五小时使用限制。我们还在为订阅用户提供速率限制重置功能,您现在可以保存并在任何时候使用它。
沟通。Opus 5.5 的沟通方式比之前的模型更自然。早期测试者发现它的写作更清晰、更容易理解,这解决了一些关于 Opus 5 的常见反馈。它将最重要的信息放在前面,其风格使其在长时间会话中成为更好的工作伙伴。正如一位早期测试者所说:“它写得就像我写的一样。”在我们自己的使用中,这使得 Opus 5.5 的工作成果更易于理解和检查——这既是一个实际好处,也是一个安全优势。
Claude Sonnet 5.5 和 Claude Haiku 5.5 将在未来几周内推出,同样带来许多在性能、效率和安全性方面的改进。
性能与成本效益
在我们的基准测试中,Claude Opus 5.5 在代理式编程、计算机使用和知识工作方面领先。话虽如此,在这些能力水平上,我们发现基准测试的差距已成为衡量现实世界差异的较不可靠指标。在我们自身的使用中,Opus 5.5 与 Claude Fable 5.1 之间的差距比这些分数所显示的要小。
Opus 5.5 Fable 5.1 Opus 5 GPT-6 Astra GPT-5.6 Sol
代理式编程 Terminal-Bench 4.0¹
代理式编程 Terminal-Bench 4.0¹ 66.4% 55.8% 52.3% 57.9% 37.3%
代理式编程 FrontierCode v1.1 (Main)
代理式编程 FrontierCode v1.1 (Main) 54.4% 50.3% 48.0% 53.3% 47.5%
代理式编程 CursorBench 4.0
代理式编程 CursorBench 4.0 57.8% 51.8% 46.6% — 41.7%
知识工作 GDPval-AA v2.1
知识工作 GDPval-AA v2.1 1846 1735 1708 1542 1588
业务流程 AutomationBench²
业务流程 AutomationBench² 40.0% 31.4% 26.9% 41.4% 28.8%
多学科推理 Humanity's Last Exam
多学科推理 Humanity's Last Exam 使用工具时 67.7% 使用工具时 65.6% 使用工具时 63.6% 使用工具时 57.2% —
代理式科学研究 Terminal-Bench-Science 0.1³
代理式科学研究 Terminal-Bench-Science 0.1³ 58.7% 52.6% 29.0% 64.6% 22.4%
计算机使用 OSWorld 2.0
计算机使用 OSWorld 2.0 部分完成 81.8% 部分完成 80.7% 部分完成 74.0% — —
视觉图表识别 Chartography
视觉图表识别 Chartography 使用工具时 89.0% 使用工具时 88.4% 使用工具时 83.4% — —
除非另有说明,所有 Claude Opus 5.5 的结果均使用最大努力的自适应思维。Terminal-Bench 4.0 的结果报告的是 Claude Opus 5.5 在 xhigh 努力水平下以及 GPT-6 Astra 在高努力水平下的表现,数据由 OpenAI 提供;这些代表了每个模型的最高得分。Claude Opus 5.5 的评估启用了其生产环境中的安全限制。当安全限制介入时,网络安全任务由 Claude Opus 4.8 完成,而生物学和前沿大语言模型开发任务则由 Claude Opus 5 完成。这可能会降低 Claude Opus 5.5 在这些基准测试中的表现。
1 Terminal-Bench 4.0:Claude Opus 5.5 的标准误差为 ±2.6 分,其他 Claude 模型的标准误差为 ±1.6–2 分。公共排行榜(每项任务 5 次试验,使用 Claude Code 框架)报告 Claude Opus 5 得分为 51.8%;我们的设置复现结果为 52.3%,在噪声范围内。GPT-6 Astra 和 GPT-5.6 Sol 的数据由 OpenAI 提供。
2 AutomationBench:AutomationBench 的结果由 Zapier 运行并报告。这些运行未使用备用模型,因此安全限制介入被视为失败——这导致得分低于 Claude Opus 5.5 在实际操作中的表现。Claude Opus 5.5 的结果来自 Zapier 在早期访问期间的自身评估。Opus 5、GPT-5.6 Sol 和 GPT-6 Astra 的结果来自 Zapier 的公共排行榜。
3 Terminal-Bench-Science 0.1:每个模型的标准误差为 ±3.5–5 分。公共排行榜(每项任务 3 次试验,使用 Claude Code 框架)报告 Claude Opus 5 得分为 30.0%;我们的设置复现结果为 29.0%,在噪声范围内。GPT-6 Astra 的数据由 OpenAI 提供。
Opus 5.5 优势非常明显的地方在于效率。它的每 token 成本低于 Opus 5,且每项任务使用的 token 更少,最终导致成本降低 40%。
定价
每百万 token 价格 Claude Opus 5.5 Claude Opus 5
缓存读取 $0.20 $0.50
输入 token $4 $5
输出 token $20 $25
缓存写入 $5 $6.25
Opus 5.5 的快速模式也可在 Claude Code 和 Claude Platform 中使用,速度最高可达 2.5 倍。其价格为每百万输入 token 8 美元,每百万输出 token 40 美元。
编程
Opus 5.5 在处理代码库范围迁移和审计等漫长且复杂的任务方面表现尤为出色。一位早期测试者使用它在不到三小时内完成了一个包含 20 万行代码的代码库的审计与修复工作,而 Opus 5 则花费了超过 20 小时,并消耗了 2.5 倍的 token 数量。在一次内部测试中,我们要求 Opus 5.5 和 Fable 5.1 将广泛用于平衡服务器间网络流量负载的 HAProxy 软件从 C 语言重写为 Rust 语言。两者的重写版本都通过了 HAProxy 自身几乎所有的回归测试,但 Opus 5.5 仅用时 9.5 小时,而 Fable 5.1 耗时 12 小时,且成本降低了 51%。
Opus 5.5 以极低的成本在代理式编码领域实现了前沿水平的成果。在 FrontierCode 的默认努力级别下,它以每项任务约 20% 的成本超越了 GPT-6 Astra。在 Terminal Bench 4.0 上,它仅以约 40% 的成本就达到了与 Astra 相当的水平,而在 CursorBench 上,它以约三分之一的成本比 GPT-5.6 Sol 高出 11 分。
代理式终端编码 代理式编码:FrontierCode 代理式编码:CursorBench
代理式终端编码 代理式编码:FrontierCode 代理式编码:CursorBench
Terminal-Bench 4.0 准确率与成本 Opus 5.5
Fable 5.1
Opus 5
GPT-6 Astra
GPT-5.6 Sol
0 10 20 30 40 50 60 70 得分 (%) 2 5 10 20 每次尝试的成本 (美元,对数刻度) 低 中 高 极高 最高
FrontierCode v1.1,主数据集 准确率与成本 Opus 5.5
Fable 5.1
Opus 5
GPT-6 Astra
GPT-5.6 Sol
35 40 45 50 55 0 得分 (%) 0.50 1 2 5 10 每项任务的成本 (美元,对数刻度) 低 中 高 极高 最高
CursorBench 4.0 准确率与成本 Opus 5.5
Fable 5.1
Opus 5
GPT-5.6 Sol
25 30 35 40 45 50 55 60 0 得分 (%) 1 2 5 10 20 每项任务的成本 (美元,对数刻度) 低 中 高 极高 最高
我们的早期测试者报告了类似的效率和智能提升:
GitHub Clio Lovable Quantium Spotify Optiver Column Kiro
GitHub Clio Lovable Quantium Spotify Optiver Column Kiro
引语 “开发者希望代理能够承担实际的软件开发工作并完成它。在我们在 GitHub Copilot CLI 和 VS Code 中的测试中,Claude Opus 5.5 使用的 token 数量和步骤数量都是我们测量中最少的之一。在 VS Code 中,它在不到一半的步骤内解决了比 Opus 5 更多的终端任务。这不仅仅是让单个任务更高效,而是让开发者的更大项目变得更具可行性。”
公司 GitHub
作者 Mario Rodriguez,首席产品官
引语 “我将一个跨越我们六个仓库的大型工程任务交给 Claude Opus 5.5,并让它整夜无人值守地运行。它在超过 18 小时内一直专注于定义我们的服务如何相互通信,并确定每个服务应如何应用这些规则。与 Opus 5 相比,它更快地达到里程碑,且需要极少返工。它的代码注释简短且有用,而不是冗长且充满散文式描述。我很难找到任何负面的评价。”
公司 Clio
作者 Sean Heintz,高级软件开发者
引语 “对于 Lovable 的构建者来说,Opus 5.5 意味着在保持相同质量的同时实现更快的构建,无论是从头开始还是处理实时应用。它一次性收集上下文,进行更少但更完整的编辑,并且不会卡在重试循环中,完成的步骤数减少三分之一到一半,同时显著减少了 token 的使用量。”
公司 Lovable
作者 Fabian Hedin,首席技术官兼联合创始人
引语 “我们在聊天、协作和 Claude Code 等我们团队使用的全部工作范围内测试了 Claude Opus 5.5。一个以前需要四天时间、38 个提示的复杂编码任务,现在仅需三个小时、11 个提示即可完成,且输出更具生产就绪性,返工更少。对于我们团队快速解决复杂问题而言,这意味着更少的迭代时间和更多的质疑时间:测试假设、对输出进行压力测试,并为我们的客户找到最佳解决方案。”
公司 Quantium
作者 Harley Barnes,AI 技术执行经理
引语 “通过 Claude Opus 5.5,我们在内部评估中看到了 token 效率的明显提升,因为我们能够以更低的成本和更快的速度完成相同的任务。”
公司 Spotify
作者 Aleksandar Mitic,高级工程师
引语:“我们在真实的工程和交易台工作中测试模型。在代理编程任务中,Claude Opus 5.5 以大约一半的交互轮次、时间和输出令牌,达到了与 Opus 5 相当的质量,将该工作负载的成本降低了 40% 到 50%。它在一家交易台的交易支持套件上创下了我们记录的最高分数,完成了此前 Claude 模型未能通过的任务,并在我们的分析任务中位列所有八个模型之首。”
公司:Optiver
作者:Noyan Tokgozoglu,全球 AI 工程负责人
引语:“Claude Opus 5.5 在委派子代理方面更加高效,并以创造性的方式检查自己的工作。设置自我验证循环变得更加容易。它在我们之前的账单中发现了以前模型遗漏的节省机会,并在代码审查中通过检查我们早期几次提交中建模错误的第三方集成的外部文档,发现了一个漏洞。”
公司:Column
作者:Mitch Fierro,工程部门
引语:“代理发出的每一次调用都是开发者所感受到的时间和成本。在真实命令行任务的一个公开基准测试中,Claude Opus 5.5 解决了比 Opus 5 更多的任务,同时调用的次数减少了约 40%,使用的令牌量减半。对于使用 Kiro 进行开发的开发者来说,这意味着无论是常规任务还是复杂挑战,代理会话都更快、更经济。Opus 5.5 即将在 Kiro 中可用。”
公司:Kiro
作者:Deepak Singh,代理 AI 副总裁
最安全的编码代理
在企业系统中使用代理的组织需要确保这些代理按预期运行,特别是当它们自主运行数小时时。Opus 5.5 拥有一个分类器,可在每次操作运行前对其进行筛选;一个可供安全团队审计的开源沙箱;以及能在漏洞合并前捕获它们的代码审查功能。
模型本身也具备更强的防御能力。在提示注入攻击方面,它在我们要测试的所有场景(包括编程、工具使用、计算机使用和网页浏览)中均持平或优于 Opus 5。在 AI 安全公司 Gray Swan 运行的基准测试中,Opus 5.5 与 Fable 5.1 并列,成为所有被测试模型中提示注入成功率最低的模型。
知识工作
Opus 5.5 是一位可靠且熟练的研究员。在一次内部测试中,我们要求 Opus 5.5、Fable 5.1 和 Opus 5 仅利用它们能在一份难以定位盈利报告的网页副本上找到的信息,撰写一份关于某公司季度业绩的报告。自动评分器将每个数据和引文与来源进行了核对。在不同的努力设置下,Opus 5.5 的 18 份报告中有 16 份达到了我们的质量门槛,而任何虚构的数据或引文都会导致失败。Fable 5.1 和 Opus 5 在任何尝试中均未达到该门槛。
它在财务分析和商业工作中也表现出色。投资公司 Walleye Capital(早期测试者)报告称,Opus 5.5 在其最低设置下基本解决了他们的评估套件;在更高设置下,其表现甚至更好,注意到了他们评估指令中的错误并进行了修正。此前没有其他模型发现过这一错误。
在另一项测试中,我们要求 Opus 5.5 和 Opus 5 分析两家虚构的人力资源软件公司之间的拟议合并。两者都在 Excel 中构建了财务模型,然后将其转化为关于该交易在其价格下是否合理的执行层演示文稿。两个模型对交易的结论相同,但 Opus 5.5 的模型更加详尽,其演示文稿也更容易阅读,而 Opus 5 的模型存在细微错误。Opus 5.5 耗时 63 分钟完成,而 Opus 5 耗时 93 分钟,且成本降低了 50%。
在知识工作评估中,Opus 5.5 在表现优于其他模型的同时,使用的令牌也更少。在 GDPval-AA v2.1(一项涵盖 44 种职业的真实工作测试)中,Opus 5.5 获得了 1846 Elo 的分数,领先于 Fable 5.1 和 Opus 5。在默认努力设置(中等)下,Opus 5.5 以每任务约五分之一的成本,在最大努力下击败了 GPT-6 Astra。它同样在其他衡量业务流程和大规模数据收集的基准测试中表现优于其他模型。
GDPval-AA v2.1 AutomationBench WANDR
GDPval-AA v2.1 AutomationBench WANDR
GDPval-AA v2.1 Elo 与成本 Opus 5.5
Fable 5.1
Opus 5
GPT-6 Astra
GPT-5.6 Sol
1200 1300 1400 1500 1600 1700 1800 0 Elo 0.20 0.50 1 2 5 10 每项任务预估成本(美元,对数刻度)低 中 高 极高 最高
AutomationBench 准确率与成本 Opus 5.5
Opus 5
GPT-6 Astra
GPT-5.6 Sol
0 10 20 30 40 通过率(%)0.50 1 2 每项任务成本(美元,对数刻度)低 中 高 极高 最高
WANDR 准确率与成本 30 40 50 60 70 0 得分(%)1 2 5 10 20 50 每次尝试成本(美元,对数刻度)低 中 高 极高 最高
我们的客户报告了类似的结果。以下是他们关于使用该模型的工作反馈:
德勤咨询 LLP Rogo LexisNexis Legal & Professional Walleye Capital Hex Thomson Reuters Labs Hebbia Viktor
德勤咨询 LLP Ro
We’re introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.
Claude Opus 5.5 is our first release since we called for pacing the frontier . It was tested before release by external evaluators, including Frontier Design and METR . On our automated behavioral audit, the most comprehensive alignment test we run, Opus 5.5 is the strongest-performing model we’ve tested to date. It also comes with the safeguards we’ve developed for our most capable models.
Here are some of the improvements you can expect from Opus 5.5:
Performance. Opus 5.5 is a major step up from Opus 5. It’s the new leading model, and early testers saw large jumps in performance on their most complex work. One tester completed a 680,000-line code migration in less than a day—work that would have taken an engineering team weeks. It’s good at finding and fixing inefficiencies in software: when we asked it to cut load times across every page of a web app, Opus 5.5 succeeded 39 of 40 times, while Opus 5 made smaller improvements that also altered the app’s behavior. A different tester had several Claude models build a game from a single prompt; Opus 5.5 scored higher than any other model on the strength of its graphics and polish.
Safety. Opus 5.5 achieves the best scores of any model to date on our automated behavioral audit, our alignment suite that tests Claude across thousands of simulated scenarios. It is much less likely than recent models to take hard-to-reverse actions or act outside the boundaries it’s been given, and it’s more resistant than Opus 5 to prompt injection. We’ve also broadened our alignment testing to cover longer tasks, impossible tasks, and scenarios modeled on real incidents, though it still has limits. Full details of our evaluation are available in the Opus 5.5 System Card .
Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1. Vetted organizations can apply today to our Life Sciences Verification Program to use Opus 5.5 for biology research. In the coming weeks we will also be expanding access to our Cyber Verification Program , and verified cybersecurity practitioners will be able to use Opus 5.5 for their work.
Cost and speed. Opus 5.5 requires less compute to serve than Opus 5, and its pricing reflects that. Our tests show that at default settings it will cost 40% less than Opus 5 on typical workloads. Input and output tokens are $4 and $20 per million, 20% less than Opus 5. Cache reads (which make up the majority of agentic and coding work costs) are $0.20 per million tokens, 60% less than Opus 5. Opus 5.5 also generates output more than 30% faster than Opus 5.
In addition to the price drop, we’re increasing five-hour usage limits on Pro, Max, Team, and seat-based Enterprise plans. We’re also providing subscription users a rate limit reset, which you can now save and use whenever you choose.
Communication. Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5. It puts the most important information up front, and its style makes it a better work partner over long sessions. As one early tester put it, “it writes the way I do.” In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one.
Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks, with many of the same improvements to performance, efficiency, and safety.
Performance and cost-effectiveness
On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.
Opus 5.5 Fable 5.1 Opus 5 GPT-6 Astra GPT-5.6 Sol
Agentic coding Terminal-Bench 4.0¹
Agentic coding Terminal-Bench 4.0¹ 66.4% 55.8% 52.3% 57.9% 37.3%
Agentic coding FrontierCode v1.1 (Main)
Agentic coding FrontierCode v1.1 (Main) 54.4% 50.3% 48.0% 53.3% 47.5%
Agentic coding CursorBench 4.0
Agentic coding CursorBench 4.0 57.8% 51.8% 46.6% — 41.7%
Knowledge work GDPval-AA v2.1
Knowledge work GDPval-AA v2.1 1846 1735 1708 1542 1588
Business workflows AutomationBench²
Business workflows AutomationBench² 40.0% 31.4% 26.9% 41.4% 28.8%
Multidisciplinary reasoning Humanity's Last Exam
Multidisciplinary reasoning Humanity's Last Exam 67.7% with tools 65.6% with tools 63.6% with tools 57.2% with tools —
Agentic scientific research Terminal-Bench-Science 0.1³
Agentic scientific research Terminal-Bench-Science 0.1³ 58.7% 52.6% 29.0% 64.6% 22.4%
Computer use OSWorld 2.0
Computer use OSWorld 2.0 81.8% partial 80.7% partial 74.0% partial — —
Visual chart recognition Chartography
Visual chart recognition Chartography 89.0% with tools 88.4% with tools 83.4% with tools — —
Unless otherwise noted, all Claude Opus 5.5 results use adaptive thinking at max effort. Terminal-Bench 4.0 results are reported for Claude Opus 5.5 at xhigh effort and GPT-6 Astra at high effort, as reported by OpenAI; these represent each model’s highest score. Claude Opus 5.5 was evaluated with its production safeguards enabled. When they intervened, cybersecurity tasks were completed by Claude Opus 4.8, and biology and frontier LLM development tasks were completed by Claude Opus 5. This likely reduces Claude Opus 5.5’s performance on these benchmarks.
1 Terminal-Bench 4.0: The standard error is ±2.6 pts for Claude Opus 5.5 and ±1.6–2 pts for the other Claude models. The public leaderboard (5 trials/task, Claude Code harness) reports Claude Opus 5 at 51.8%; our setup reproduces it at 52.3%, within noise. GPT-6 Astra and GPT-5.6 Sol figures are as reported by OpenAI.
2 AutomationBench: AutomationBench results were run and reported by Zapier. These runs were performed without fallback models, so safeguard interventions were considered failures—this resulted in a lower score than Claude Opus 5.5 would achieve in practice. Claude Opus 5.5 results come from Zapier’s own evaluation during early access. Results for Opus 5, GPT-5.6 Sol, and GPT-6 Astra come from Zapier’s public leaderboard.
3 Terminal-Bench-Science 0.1: The standard error is ±3.5–5 pts per model. The public leaderboard (3 trials/task, Claude Code harness) reports Claude Opus 5 at 30.0%; our setup reproduces it at 29.0%, within noise. The GPT-6 Astra figure is as reported by OpenAI.
Where Opus 5.5’s advantage is very clear is efficiency. It costs less per token than Opus 5 and uses fewer tokens per task, which nets out to a 40% drop in costs.
Pricing
Prices per 1M tokens Claude Opus 5.5 Claude Opus 5
Cache reads $0.20 $0.50
Input tokens $4 $5
Output tokens $20 $25
Cache writes $5 $6.25
Fast mode for Opus 5.5 is also available in Claude Code and the Claude Platform with up to 2.5x speed. It costs $8 per million input tokens and $40 per million output tokens.
Coding
Opus 5.5 is particularly good at long and sprawling jobs like codebase-wide migrations and audits. An early tester used it to audit and fix a 200,000-line codebase in under three hours, where Opus 5 took over 20 hours and used 2.5x as many tokens. In an internal test, we asked Opus 5.5 and Fable 5.1 to translate HAProxy, widely used software that balances web traffic loads across servers, from C into Rust. Both rewrites passed nearly all of HAProxy’s own regression tests, but Opus 5.5 finished in 9.5 hours compared to 12 for Fable 5.1, and cost 51% less.
Opus 5.5 delivers frontier results on agentic coding at a fraction of the cost. At its default effort level on FrontierCode, it beats GPT-6 Astra at roughly 20% of the cost per task. On Terminal Bench 4.0, it matches Astra for about 40% of the cost, while on CursorBench it beats GPT-5.6 Sol by 11 points for about a third of the cost.
Agentic terminal coding Agentic coding: FrontierCode Agentic coding: CursorBench
Agentic terminal coding Agentic coding: FrontierCode Agentic coding: CursorBench
Terminal-Bench 4.0 Accuracy vs Cost Opus 5.5
Fable 5.1
Opus 5
GPT-6 Astra
GPT-5.6 Sol
0 10 20 30 40 50 60 70 Score (%) 2 5 10 20 Cost per attempt (USD, log scale) low med high xhigh max
FrontierCode v1.1, main set Accuracy vs Cost Opus 5.5
Fable 5.1
Opus 5
GPT-6 Astra
GPT-5.6 Sol
35 40 45 50 55 0 Score (%) 0.50 1 2 5 10 Cost per task (USD, log scale) low med high xhigh max
CursorBench 4.0 Accuracy vs Cost Opus 5.5
Fable 5.1
Opus 5
GPT-5.6 Sol
25 30 35 40 45 50 55 60 0 Score (%) 1 2 5 10 20 Cost per task (USD, log scale) low med high xhigh max
Our early testers reported similar efficiency and intelligence gains:
GitHub Clio Lovable Quantium Spotify Optiver Column Kiro
GitHub Clio Lovable Quantium Spotify Optiver Column Kiro
Quote “Developers want agents that can take on real software work and finish it. In our testing across GitHub Copilot CLI and VS Code, Claude Opus 5.5 used among the fewest tokens and steps we measured. In VS Code, it solved more terminal tasks than Opus 5 in less than half the steps. More than making individual tasks more efficient, it’s making developers’ bigger projects more achievable.”
Company GitHub
Author Mario Rodriguez, Chief Product Officer
Quote “I handed Claude Opus 5.5 a large engineering task across six of our repositories and let it run overnight, unattended. It stayed on task for over 18 hours defining how our services talk to each other and working out how each one should apply that. Compared with Opus 5, it hit milestones faster and required minimal reworking. Its code comments were short and useful instead of long and prose-heavy. I’m struggling to find anything negative to say.”
Company Clio
Author Sean Heintz, Staff Software Developer
Quote “For Lovable builders, Opus 5.5 means faster builds with the same quality, whether you’re starting from scratch or working on a live app. It gathers context once, makes fewer and more complete edits, and doesn’t get stuck retrying, finishing in a third to half fewer steps and using significantly fewer tokens along the way.”
Company Lovable
Author Fabian Hedin, CTO and Co-founder
Quote “We tested Claude Opus 5.5 across Chat, Cowork, and Claude Code, the full range of how our teams work. A complex coding task that previously took 38 prompts over four days came in at 11 prompts over three hours, with more production-ready outputs and less rework. For our teams solving complex problems at pace, that means less time iterating and more time interrogating: testing assumptions, pressure-testing outputs, and landing on the best solution for our clients.”
Company Quantium
Author Harley Barnes, Executive Manager, AI Technology
Quote “With Claude Opus 5.5, we’ve seen a clear improvement in token efficiency across our internal evaluations, as we’ve been able to complete the same tasks both cheaper and faster.”
Company Spotify
Author Aleksandar Mitic, Senior Engineer
Quote “We test models on real engineering and trading-desk work. On our agentic coding tasks, Claude Opus 5.5 matched Opus 5’s quality in about half the turns, time and output tokens, cutting the cost of that workload by 40 to 50%. It posted the highest score we’ve recorded on one desk’s trading-support suite, passing tasks earlier Claude models had failed, and topped all eight models on our analysis task.”
Company Optiver
Author Noyan Tokgozoglu, Global Head of AI Engineering
Quote “Claude Opus 5.5 delegates to subagents far more effectively and checks its own work in creative ways. Self-verification loops feel easier to set up. It found savings opportunities in our cloud bill that previous models had missed, and in code review it caught a bug by checking external docs for a third-party integration we’d modeled wrong several commits earlier.”
Company Column
Author Mitch Fierro, Engineering
Quote “Every call an agent makes is time and cost a developer feels. On a public benchmark of real command-line tasks, Claude Opus 5.5 solved more than Opus 5 while making about 40% fewer calls and using half the tokens. For developers building with Kiro, that means faster, more affordable agent sessions for routine tasks and complex challenges alike. Opus 5.5 will soon be available in Kiro.”
Company Kiro
Author Deepak Singh, VP of Agentic AI
The most secure coding agent
Enterprises that use agents within their systems need to know that those agents are operating as intended, particularly when they run autonomously for many hours. Opus 5.5 has a classifier that screens every action before it runs, an open-source sandbox that security teams can audit, and code review that catches vulnerabilities before they merge.
The model itself also has stronger defenses. On prompt injection attacks, it matches or beats Opus 5 in every setting we tested, including coding, tool use, computer use, and web browsing. On a benchmark run by the AI security firm Gray Swan, Opus 5.5 ties Fable 5.1 for the lowest prompt injection success rate of any model tested.
Knowledge work
Opus 5.5 is a reliable and adept researcher. In one internal test, we asked Opus 5.5, Fable 5.1, and Opus 5 to write a report on a company’s quarterly performance using only the information it could find on a copy of the web where the earnings release was hard to locate. An automated grader checked every figure and quote against sources. Across different effort settings, 16 out of 18 of Opus 5.5’s reports cleared our quality bar, where any invented figure or quote would have failed. Neither Fable 5.1 nor Opus 5 cleared that bar in any attempt.
It’s also strong in financial analysis and business work. Walleye Capital, an investment firm and early tester, reported that Opus 5.5 largely solved their evaluation suite on its lowest setting; on higher settings, it performed even better, noticing an error in their evaluation instructions and correcting for it. No other model had caught this error before.
In another test, we tasked both Opus 5.5 and Opus 5 with analyzing a proposed merger between two fictional HR software companies. Each built a financial model in Excel, then turned it into an executive presentation on whether the deal made sense at its price. Both models reached the same conclusions about the deal, but Opus 5.5’s model was more thorough and its presentation easier to read, while Opus 5’s had minor errors. Opus 5.5 finished in 63 minutes compared to 93 for Opus 5, and cost 50% less to produce.
On knowledge work evaluations, Opus 5.5 outperforms other models while also using fewer tokens. On GDPval-AA v2.1, a test of real-world work across 44 occupations, Opus 5.5 scores 1846 Elo, ahead of Fable 5.1 and Opus 5. At default effort (medium), Opus 5.5 beats GPT-6 Astra at max effort for about a fifth of the cost per task. It likewise outperformed other models on benchmarks measuring business workflows and large-scale data collection.
GDPval-AA v2.1 AutomationBench WANDR
GDPval-AA v2.1 AutomationBench WANDR
GDPval-AA v2.1 Elo vs Cost Opus 5.5
Fable 5.1
Opus 5
GPT-6 Astra
GPT-5.6 Sol
1200 1300 1400 1500 1600 1700 1800 0 Elo 0.20 0.50 1 2 5 10 Estimated cost per task (USD, log scale) low med high xhigh max
AutomationBench Accuracy vs Cost Opus 5.5
Opus 5
GPT-6 Astra
GPT-5.6 Sol
0 10 20 30 40 Pass rate (%) 0.50 1 2 Cost per task (USD, log scale) low med high xhigh max
WANDR Accuracy vs Cost 30 40 50 60 70 0 Score (%) 1 2 5 10 20 50 Cost per attempt (USD, log scale) low med high xhigh max
Our customers have reported similar results. Here’s what they told us about working with the model:
Deloitte Consulting LLP Rogo LexisNexis Legal & Professional Walleye Capital Hex Thomson Reuters Labs Hebbia Viktor
Deloitte Consulting LLP Ro
首次收录 · 2026-09-23 · 10.17 分