任务成本几何?
Claude Code 中的一个任务是一个循环。模型阅读对话内容,调用工具,读取结果,然后再次循环,直到任务完成。每完成一次循环即产生一次请求。决定循环成本的有四个因素。
轮次(Turns)。每一轮都会重新发送迄今为止的对话内容。轮次越少,处理的输入量就越少。
缓存读取。每一轮重新发送的内容中,大部分是模型在前一轮已经看过的文本。这部分按“缓存读取”计费,价格仅为输入价格的极小一部分。
输出令牌类型。最昂贵的令牌是输出令牌,其价格是输入价格的五倍。“思考”过程被计为输出,因此,在得出答案过程中推理较少的模型成本更低。
模型。每种模型都有各自的价格,详见定价页面,因此你选择的模型决定了每个令牌的价格。
我们的示例使用 Opus 5.5 API 列表价格:每百万输入令牌 4 美元,每百万输出令牌 20 美元,以及每百万缓存读取 0.20 美元。与下文中的计算器一样,这些示例将缓存输入按读取价格计费,其余部分按输入价格计费,并省略了缓存写入费用。令牌数量仅为示意。
轮次
假设一个任务以 2 万令牌的上下文开始,随着模型阅读文件和工具结果,上下文增长至 12 万。在 40 轮的情况下,平均每轮发送约 7 万令牌。该任务的输入令牌总数约为 280 万,尽管对话内容从未超过 12 万。如果其中 90% 从缓存中读取,输入成本约为 1.62 美元。如果在 25 轮内完成同一任务,处理的令牌数约为 175 万,输入成本约为 1.02 美元。
一轮的成本高于其新增的令牌成本,因为它会重新发送之前的所有内容。因此,最便宜的一轮是你不需要的那一轮。
减少轮次的一个习惯是给予模型检查其工作成果的方法。例如,运行测试、构建项目或调用端点的脚本。能够自行检查工作成果的模型能更早发现错误。
一次性收集所需信息并批量调用工具的模型,也能减少重新发送的次数。
缓存读取
如果 280 万输入令牌中没有来自缓存的部分,成本为 11.20 美元。在命中率为 90% 时,成本为 1.62 美元;在命中率为 96% 时,成本约为 0.99 美元。没有其他设置能像这样大幅影响输入成本。持续的会话自然会保持较高的命中率。我在本文稍后部分会介绍一些避免破坏缓存的操作。
输出令牌
在 Opus 5.5 上,一个输出令牌的成本是缓存读取的 100 倍。典型任务的 6 万输出令牌成本为 1.20 美元,相当于从缓存中读取 600 万令牌的开销。输出包括“思考”过程。你需要为所有输出付费,即使 Claude Code 仅向你展示摘要。这就是为什么主要改变模型思考量的“努力程度(effort)”设置会对账单产生如此大的影响。
模型
拥有更便宜缓存读取的模型主要有助于长会话。拥有更便宜输出令牌的模型主要有助于需要大量推理的任务。
Opus 5.5 中发生了什么变化?
有两点发生了变化:价格,以及模型完成的工作量。
每一项价格都降低了。输入和输出令牌比 Opus 5 便宜 20%。缓存读取便宜了 60%。输入价格下降,其读取费率也随之下降,从输入价格的十分之一降至二十分之一。图 A 比较了两种模型每百万令牌的价格。这些是 API 列表价格。在 Pro、Max 或 Team 套餐中,较低的 Opus 5.5 价格会反映在你的额度限制中,包括缓存上下文,因此你的额度比在 Opus 5 上多出约 25%。缓存读取的额外折扣是 API 价格的调整。
此外,Pro、Max、Team 以及基于席位的 Enterprise 计划的五小时限制已上调,符合条件的订阅者获得一个可自主决定使用时间的限制重置机会。你可以在网页版或 Claude Desktop 的“设置 > 用量”中找到它,而不是在终端中的 Claude Code 里。该重置适用于你的整个账户,包括 Claude Code。
图 A API 每百万令牌列表价格。
对于 API 密钥而言,缓存读取价格的降低对 Claude Code 影响最大。一个长期的智能体会话将其大部分输入花费在缓存读取上。在下文图 B 所示的会话中,缓存费用从 1.00 美元降至 0.40 美元,这是账单中最大的降幅。
节省多少费用取决于你工作的形态。一个以缓存读取为主的会话,输入端最高可节省 60%。一个没有缓存、短问题配长回答的会话,由于输出占主导地位,最高可节省 20%。大多数 Claude Code 任务介于两者之间。下方的计算器可以显示你的任务处于什么位置。
你可能已经注意到,Opus 5.5 的运行成本比 Opus 5 低 40%。这是我们对按令牌计费、默认设置下的典型工作负载的估算。它假设在较低价格的基础上,Opus 5.5 在其默认的中等水平下每个任务使用的令牌更少,因此这并非意味着单个令牌的单价直接削减了 40%。令牌价格如 Fig A 所示,Fig B 展示了仅价格变化带来的影响。
Opus 5.5 在回答中可以使用更多令牌,因为它总是先思考再回复。我们预计人们能在 Opus 5.5 上完成更多工作,但这因任务而异,所以请根据你的实际工作进行测量。Fig C 比较了两个模型每个任务的成本。这部分内容对你的工作的依赖程度远高于对价格的依赖。
对于范围明确的任务,两个模型完成的轮数大致相同,你得到的全部收益就是价格削减。差距在开放式任务中应该最大,因为模型可能会在错误的思路上花费很多轮次。没有哪个单一数字适用于所有代码库,所以请进行测量(见最后一节)。
长会话以报告结束。Opus 5.5 会在长会话结束时提供它所做的更改、发现的内容以及它需要你提供的信息。这也能省钱,因为当你能够看到发生了什么时,你重新运行会话的频率会降低。
最大化会话价值的技巧
Opus 5.5 较低的价格使得每个令牌的成本更低。你如何运行一个会话决定了你使用多少令牌,以下步骤会有所帮助。
在更改模型之前提高努力程度
“努力程度”(Effort)设定了模型每轮花费令牌的总体倾向:包括其思考、撰写的文本以及工具调用。在较低的努力程度下,它进行的工具调用更少且更简短。Opus 5.5 有四个级别(低、中、高和超高),加上针对单个会话的“最大”(max)。选择下方的一个级别以查看何时使用它以及设置它的命令。
低:机械性工作
中:日常事务
高:如果停滞不前
超高:难题
最大:单次会话
中等:适用于范围明确的日常工作
具有明确范围的日常工作。它在每轮中的思考比“高”少,因此每轮成本更低。
当任务范围明确时从这里开始。
/effort medium 复制
/effort status 复制
/model 复制
/usage 复制
级别从每轮最少思考到最多思考排列。
Claude Code 为每个模型设置了一个默认级别,/effort status 会显示你的当前级别。对于范围明确、日常的工作尝试“中等”。当“中等”停滞不前时,尝试“高”。它在每轮上花费的比“中等”多,但比切换到更大的模型少。对于机械性工作(如重命名或在多个文件中应用已知模式)使用“低”。
在 Opus 5.5 上,默认级别是“中等”,比 Opus 5 的默认级别“高”低一级。不同模型在同一级别下的思考量并不相同。在给定级别下,Opus 5.5 每轮的思考量多于 Opus 5,尤其是在“超高”和“最大”级别。因此,不要沿用你为 Opus 5 选择的级别。从“中等”开始,并将“超高”和“最大”保留在你已测量到收益的工作上。
关于努力程度定价的一个粗略思考方式:假设“高”在一个任务中增加了 20K 个思考令牌。在 Opus 5.5 上这相当于 0.40 美元。一个包含十轮重试、100K 缓存上下文和总共 10K 输出令牌的循环,成本大致相同。因此,“高”在一个能节省一次重试的任务中是物有所值的。在一个“中等”本可以一次性完成的任务中,则是浪费。
当“中等”只能修复一层时
你需要更多努力程度的最明显迹象是一个只停留在某一层的修复。
假设 API 处理程序中的某个字段被重命名。在“中等”级别下,模型更新了处理程序,处理程序的测试通过,但客户端仍在发送旧字段。它完成了被要求做的事。只是它没有读取得足够远以找到第二个调用者。在“高”级别下,它在撰写代码之前会花费更多轮次来阅读调用点,并在一次传递中更改这两个层级。
检查可以发现相同的错误。如果模型能够运行通过客户端的测试,那么旧字段在写入时的那个回合就会失败,此时处于中等(medium)状态。因此,在你增加思考深度之前,先检查模型是否有自检的方法。一次测试运行消耗一个回合及其输出。更多的思考会添加到每一个回合中。
如果提升思考层级并添加检查无效,则切换到更大的模型。
在会话中途更改思考层级
在 Claude Code 中,运行 /effort 命令并指定层级,例如 /effort high。/effort status 打印当前层级。你可以在任务中途更改它,新层级将应用于下一个请求。
在使用 API 密钥或 Claude 订阅的 Opus 5.5 上,更改思考层级会保留缓存。你可以为某个困难步骤提高层级,然后再降低,而无需重写对话。在 Amazon Bedrock、Google Cloud 的 Agent Platform 或 Claude 应用网关上,更改思考层级仍会清除缓存的对话,且下一个请求需支付整个对话的缓存写入费用。对于 Opus 5.5,思考功能始终开启,因此没有可更改的思考设置。
为你的工作选择合适的模型
模型选择决定了会话中每个令牌的价格,因此它对账单的影响比思考层级更大。它的影响也更深远。每个继承主模型的子代理也会继承其价格。大多数情况下需要三种模型:一种用于查找的小模型,一种用于你密切监督的工作的 Opus 5.5,以及一种用于最困难任务的大模型。
将 Opus 5.5 作为日常主力
将你监督的工作交给 Opus 5.5 处理:跨几个文件的功能开发、调试以及带有后续编辑的代码审查。你阅读它的工作成果并在其偏离时介入,从而使循环保持简短。
升级到 Fable 5.1
当结果的重要性超过令牌价格时,升级到 Fable 5.1。例如,你无法监督的长运行任务、代码库中不存在现有模式的问题,以及协调多个子代理的大规模更改。不要等到第三次失败。如果 Opus 5.5 在 xhigh 层级上两次遇到相同的问题,请切换模型,并在问题解决后切回。对于交互式工作,Opus 5.5 更为合适,因为它的延迟更低且成本更少。
Fable 5.1 的定价为每百万输入令牌 10 美元,每百万输出令牌 50 美元,是 Opus 5.5 价格的 2.5 倍。其缓存读取费用为每百万 0.25 美元,仅为 Opus 5.5 费率的 1.25 倍,因为它们按 Opus 5.5 输入价格的 0.025 倍计费。因此,差距在长且大量使用缓存的运行中最小,而在写入量大的任务中最大。
在自然断点处切换。缓存属于前一个模型,因此预计在新模型上的第一个回合需支付整个对话的写入费用。先运行 /compact,或开始一个新的会话并附上简短的书面计划,以使该回合更小。运行 /model 命令并指定别名或模型名称以进行切换。/model 还会将你的选择保存为新会话的默认设置,因此在完成困难部分后请切回。
降级用于查找
降级到 Sonnet 或 Haiku 用于查找,而非编写代码:搜索和摘要的子代理、阅读日志和测试输出,以及“此处在何处定义”的问题。对于跨多个文件的机械性编辑,保留 Opus 5.5 并将思考层级设为低(low)。编辑操作仍由编写你其余代码的模型执行,每个回合的成本更低。
要将子代理置于较小的模型上,请在其定义中设置 model: haiku 或 model: sonnet。要将所有子代理置于同一模型上,请设置 CLAUDE_CODE_SUBAGENT_MODEL 环境变量。子代理定义中指定的模型名称会覆盖该变量。没有模型设置的子代理将在你的主模型上运行,除非设置了该变量。
每个子代理都在其自己的上下文窗口中运行,并返回摘要,因此其文件读取操作不会进入你的主对话。它仍需为其自身的令牌付费,因此模型设置决定了这笔支出的成本。
智能体团队(Agent teams)是一项实验性功能,它会放大这一效应。每位团队成员都是独立的 Claude Code 实例,拥有各自独立的上下文窗口,并且会持续消耗令牌直到退出。我们的成本文档指出,当团队成员在计划模式(plan mode)下运行时,一个团队的令牌消耗量约为标准会话的七倍。请保持团队规模小巧,确保每个任务自包含,并在其部分工作完成后关闭团队成员。
权衡之处在于:如果一个小模型误读了搜索结果,会导致主模型去查找错误的文件,而主模型则需为这次绕道买单。请将小模型用于那些错误易于察觉的工作,例如查找文件、运行测试和读取日志。
将判断性决策留给主模型。opusplan 别名以另一种方式分配工作:Opus 在计划模式下进行规划,Sonnet 负责执行该计划。这将代码编辑任务交给了 Sonnet,与上述建议相反。在将其设为默认设置之前,请先在你的实际任务中进行测试。
迁移时检查你的提示词
为旧模型编写的指令可能导致 Opus 5.5 生成更多内容并重复工具调用。在 Claude Code 中运行 /claude-api prompt-audit 以检查你的 Claude Code 设置(如技能和 CLAUDE.md 文件)是否存在这些提示词反模式。它还会检查你在 Claude Platform 上构建的应用程序的代码。
我们在从 Opus 4.8 迁移到 Opus 5.5 的过程中对此进行了测试,使用的是一个包含 44 张工单的内部客户支持基准测试,其提示词中包含多种此类模式。转向 Opus 5.5 以较低的成本将基准测试的成本降低了约 18%。运行 prompt-audit 进一步将其降低了 9%,使其比 Opus 4.8 的起始点低约 25%。审计移除了那些导致模型生成更多内容并重复工具调用的仪式性指令:强制性的六步程序、草稿纸规则、双重验证规则以及相互矛盾的指令。
该结果来自单一基准测试,因此请将其视为示例而非预期数值。运行审计,然后在实际任务前后比较 /usage(参见“自行测量”)。
缓存与压缩
Claude Code 为你处理缓存和压缩。你如何运行会话决定了它们能节省多少。
缓存的工作原理
Claude Code 会缓存请求中重复的部分,例如系统提示词、工具定义以及迄今为止的对话内容。
在 Opus 5.5 上,缓存读取的成本仅为新输入令牌的 5%。写入缓存的成本高于新读取成本,根据当前的定价,对于五分钟缓存是输入价格的 1.25 倍,对于一小时缓存则是输入价格的两倍。每次命中都会免费重置其生命周期。
在 Claude Code 中,生命周期取决于你的付费方式。在 Claude 订阅中,它为一小时。在 API 密钥或云提供商上,默认情况下为五分钟,一旦开始使用用量积分,订阅也会降至五分钟。
在 120K 令牌的上下文下,Opus 5.5 的五分钟写入成本约为 0.60 美元,读取成本约为 0.02 美元。一次写入的成本相当于 25 次读取。在 API 密钥上,六分钟的咖啡休息时间会将下一次 0.02 美元的读取变为 0.60 美元的写入。同样大小的一次一小时写入成本约为 0.96 美元,而在 API 上,你可以支付这一溢价来填补你一天中的空档期。
会话形状与命中率
缓存存储的是前缀,因此它只能重用从开头开始与之前请求匹配的那部分请求内容。
稳定的会话会在每一轮对话中追加到对话末尾,并保持高命中率。任何改变请求早期部分的操作都会降低命中率。更改工具定义会清除整个缓存,而更改系统提示词会从该点开始清除缓存,这几乎涵盖了所有内容。
在实践中,出现以下情况时预计会发生缓存写入:
你暂停的时间超过了缓存生命周期;
你在 Amazon Bedrock、Google Cloud 的 Agent Platform 或网关上更改了努力级别(effort),因为该级别是缓存匹配的一部分;
你在对话中首次开启快速模式(fast mode),这会改变缓存匹配的部分内容;
你连接或断开 MCP 服务器,这可能会改变每个请求启动时加载的内容;
你切换模型,因为新模型从空缓存开始;以及
对话被压缩,这会重写缓存所匹配的历史记录。
因此,请在会话开始时设置这些配置,并在会话运行过程中保持不变。
为什么长会话每次交互的成本更高
E
WHAT DOES A TASK COST?
A task in Claude Code is a loop. The model reads the conversation, calls a tool, reads the result, and goes round again until it's done. Each trip round the loop is one request. Four things set what the loop costs.
Turns. Every turn resends the conversation so far. Fewer turns means less input processed.
Cache reads. Most of what a turn resends is text the model saw on the previous turn. It's billed as a cache read, at a small fraction of the input price.
Output token type. The most expensive tokens, at five times the input price. Thinking is billed as output, so a model that reasons less on the way to the answer costs less.
Model. Each model has its own prices, listed on the pricing page, so the model you pick sets the price of every token.
Our examples use Opus 5.5 API list prices: $4 per million input tokens, $20 per million output tokens and $0.20 per million cache reads. Like the calculator further down, the examples bill cached input at the read price and everything else at the input price, and leave out cache writes. The token counts are illustrations.
Turns
Letâs say a task starts with 20K tokens of context and grows to 120K as the model reads files and tool results. At 40 turns, the average turn sends about 70K tokens. That's about 2.8M input tokens for the task, though the conversation never grew past 120K. With 90% read from cache, the input costs about $1.62. The same task in 25 turns processes about 1.75M tokens and costs about $1.02 in input.
A turn costs more than the tokens it adds, because it resends everything before it. So the cheapest turn is the one you don't need.
One habit that can cut turns is giving the model a way to check its work. For example, a test to run, a build, or a script that calls the endpoint. A model that can check its own work finds its mistakes earlier.
A model that gathers what it needs in one pass, and batches its tool calls, pays the resend fewer times too.
Cache reads
The same 2.8M input tokens cost $11.20 if none come from cache. At a 90% hit rate they cost $1.62, and at 96% about $0.99. No other setting moves input cost this much. A steady session keeps a high hit rate on its own. I cover some actions to avoid breaking your cache later in this post.
Output tokens
On Opus 5.5, an output token costs 100 times a cache read. The 60K output tokens of a typical task cost $1.20, the same as reading 6M tokens from cache. Output includes thinking. You pay for all of it, even when Claude Code only shows you a summary. That's why effort, which mostly changes how much the model thinks, moves the bill so much.
Model
A model with cheaper cache reads mostly helps long sessions. One with cheaper output mostly helps tasks that need a lot of reasoning.
WHAT CHANGED IN OPUS 5.5
Two things changed: the price, and how much work the model does.
Every price line is lower. Input and output tokens are 20% cheaper than on Opus 5. Cache reads are 60% cheaper. The input price falls, and the read rate falls with it, from a tenth of the input price to a twentieth. Fig A compares the two models per million tokens. These are API list prices. On a Pro, Max or Team plan, the lower Opus 5.5 price is passed on to your limits, including cached context, so they go about 25% further than on Opus 5. The extra cut on cache reads is an API price change.
Separately, five-hour limits went up on Pro, Max, Team and seat-based Enterprise plans, and eligible subscribers get a limit reset to use when they choose. You'll find it under Settings > Usage on the web or in Claude Desktop, not in Claude Code in your terminal. A reset applies across your account, Claude Code included.
FIG A API list price per million tokens.
On an API key, the cache-read price cut matters most for Claude Code. A long agentic session spends most of its input on cache reads. In the session priced in Fig B below, the cache line falls from $1.00 to $0.40, the largest drop on the receipt.
How much you save depends on the shape of your work. A session that is mostly cache reads can save up to 60% on input. A short question with no cache and a long answer can save up to 20%, because output dominates it. Most Claude Code tasks sit between the two. The calculator below shows where yours sits.
You may have seen that Opus 5.5 costs 40% less to run than Opus 5. That's our estimate for typical workloads billed by token, at default settings. It assumes Opus 5.5 uses fewer tokens per task at its medium default, on top of the lower prices, so it isn't a 40% cut to the price of a token. The token prices are the ones in Fig A, and Fig B shows what the price change does on its own.
Opus 5.5 can use more tokens on an answer, because it always thinks before it replies. We expect people to get more done on Opus 5.5, but it varies by task, so measure it on your own work. Fig C compares cost per task on the two models. This part depends on your work far more than the price does.
On a well-scoped task, both models finish in about the same number of turns, and the price cut is all you get. The gap should be biggest on open-ended tasks, where a model can spend many turns on the wrong idea. No single number holds for every codebase, so measure it (see the last section).
Long runs end with a report. Opus 5.5 closes a long run with what it changed, what it found, and what it needs from you. That can save money too, because you rerun a session less often when you can see what happened.
TIPS FOR MAXIMIZING THE VALUE OF YOUR SESSION
Opus 5.5âs lower price makes each token cost less. How you run a session decides how many tokens you use, and these steps help.
Raise effort before you change models
Effort sets a general disposition for how many tokens the model spends on each turn: its thinking, the text it writes, and its tool calls. At lower effort it makes fewer tool calls and keeps them shorter. Opus 5.5 has four levels (low, medium, high and xhigh), plus max for a single session. Pick a level below to see when to use it and the command that sets it.
low mechanical medium everyday high if it stalls xhigh hard problems max one session
Medium: for well-scoped daily work
Day-to-day work with a clear scope. It thinks less per turn than high, so each turn costs less.
Start here when the task is well scoped.
/effort medium Copy
/effort status Copy
/model Copy
/usage Copy
Levels run from least to most thinking per turn.
Claude Code sets a default level for each model, and /effort status shows yours. Try medium for well-scoped, day-to-day work. When medium stalls, try high . It spends more per turn than medium, but less than moving to a bigger model. Use low for mechanical work, like renames or applying a known pattern across files.
On Opus 5.5 the default is medium, one level below Opus 5's default of high. Levels don't mean the same amount of thinking on every model. At a given level, Opus 5.5 thinks more per turn than Opus 5, most of all at xhigh and max. So don't carry over a level you chose for Opus 5. Start at medium, and keep xhigh and max for work where you've measured a gain.
A rough way to think about effort pricing: say high adds 20K thinking tokens across a task. On Opus 5.5 that's $0.40. A retry loop of ten turns at 100K of cached context, with 10K output tokens in total, costs about the same. So high pays for itself on a task where it saves one retry. On a task medium would have finished the first time, it's wasted.
When medium fixes one layer
The clearest sign you need more effort is a fix that stops at one layer.
Say a field is renamed in an API handler. At medium, the model updates the handler, the handler's tests pass, and the client still sends the old field. It did what it was asked. It just didn't read far enough to find the second caller. At high, it spends more turns reading call sites before it writes, and it changes both layers in one pass.
A check can catch the same bug. If the model can run a test that goes through the client, the old field fails that test on the turn it was written, at medium. So before you raise effort, check whether the model has a way to check its work. A test run costs one turn and its output. More effort adds thinking to every turn.
If upgrading effort levels and adding checks doesnât work, then switch to a bigger model.
Changing effort mid-session
In Claude Code, run /effort with a level, for example /effort high. /effort status prints the current level. You can change it mid-task, and the new level applies to the next request.
On Opus 5.5 with an API key or a Claude subscription, changing effort keeps the cache. You can raise it for one hard step and lower it again without rewriting the conversation. On Amazon Bedrock, Google Cloud's Agent Platform or a Claude apps gateway, a change of effort still clears the cached conversation, and the next request pays the cache-write price on all of it. Thinking is always on for Opus 5.5, so there's no thinking setting to change.
Choose the right model for your work
Model choice sets the price of every token in a session, so it moves the bill more than effort does. It also reaches further. Every subagent that inherits the main model inherits its price too. Most days need three models: a small one for lookups, Opus 5.5 for work you supervise closely, and a bigger one for the hardest tasks.
Opus 5.5 as the daily driver
Use Opus 5.5 for work you supervise: feature work across a few files, debugging, and code review with follow-up edits. You read what it does and step in when it drifts, so the loop stays short.
Moving up to Fable 5.1
Move up to Fable 5.1 when the result matters more than the token price. For example long runs you won't supervise, problems with no existing pattern in the codebase, and large changes that coordinate many subagents. Don't wait for a third failure. If Opus 5.5 on xhigh hits the same problem twice, switch, and switch back once it's solved. For interactive work, Opus 5.5 is a better fit as it has lower latency and costs less.
Fable 5.1 lists at $10 per million input tokens and $50 per million output, two and a half times the Opus 5.5 price. Its cache reads cost $0.25 per million, only 1.25 times the Opus 5.5 rate, because they bill at 0.025 times its input price. So the gap is smallest on a long, cache-heavy run, and largest on a task that writes a lot.
Switch at a natural break. The cache belongs to the previous model, so expect the first turn on the new model to pay the write price on the whole conversation. Run /compact first, or start a fresh session with a short written plan, to make that turn smaller. Run /model with an alias or a model name to switch. /model also saves your choice as the default for new sessions, so switch back when the hard part is done.
Moving down for lookups
Move down to Sonnet or Haiku for lookups, not for writing code: subagents that search and summarize, reading logs and test output, and "where is this defined" questions. For a mechanical edit across many files, keep Opus 5.5 and set effort to low. The edit stays on the model that writes the rest of your code, at a lower cost per turn.
To put a subagent on a smaller model, set model: haiku or model: sonnet in its definition. To put every subagent on one model, set the CLAUDE_CODE_SUBAGENT_MODEL environment variable. A model named in a subagent's definition overrides the variable. A subagent with no model setting runs on your main model, unless the variable is set.
Each subagent runs in its own context window and hands back a summary, so its file reads stay out of your main conversation. It still pays for its own tokens, so the model setting decides what that spend costs.
Agent teams, an experimental feature, multiply this. Each teammate is a separate Claude Code instance with its own context window, and it keeps using tokens until it exits. Our costs docs put a team at about seven times the tokens of a standard session when teammates run in plan mode. Keep teams small, keep each task self-contained, and shut teammates down when their part is done.
The tradeoff: a small model that misreads a search result sends the main model after the wrong file, and the main model pays for the detour. Keep the small model on work where a mistake is cheap to spot, like finding files, running tests and reading logs.
Keep judgment calls on the main model. The opusplan alias splits the work a different way: Opus plans in plan mode, and Sonnet carries out the plan. That puts the code edits on Sonnet, the opposite of the advice above. Measure it on your own tasks before you make it a default.
Check your prompts when you migrate
Instructions written for an older model can make Opus 5.5 write more and repeat tool calls. Run /claude-api prompt-audit in Claude Code to check your Claude Code setup, such as your skills and CLAUDE.md file, for these prompting anti-patterns . It also checks the code of an app you build on the Claude Platform.
We tested this on a migration from Opus 4.8 to Opus 5.5, using an internal customer support benchmark of 44 tickets whose prompt had several of these patterns. The move to Opus 5.5, at low effort, cut the benchmark's cost by about 18%. Running prompt-audit cut it by a further 9%, to about 25% below the Opus 4.8 starting point. The audit removed ritual instructions that made the model write more and repeat tool calls: a mandatory six-step procedure, a scratchpad rule, a verify-twice rule, and instructions that contradicted each other.
That result comes from one benchmark, so treat it as an example rather than a number to expect. Run the audit, then compare /usage on a real task before and after (see Measure it yourself).
Caching and compaction
Claude Code handles caching and compaction for you. How you run a session decides how much they save.
How the cache works
Claude Code caches the parts of a request that repeat, such as the system prompt, tool definitions, and the conversation so far.
On Opus 5.5 a cached read costs 5% of a fresh input token. Writing to the cache costs more than a fresh read, at 1.25 times the input price for a five-minute cache and twice the input price for a one-hour cache, on today's pricing. Each hit resets the lifetime at no charge.
In Claude Code the lifetime depends on how you pay. On a Claude subscription it's an hour. On an API key or a cloud provider it's five minutes by default, and a subscription drops to five minutes once it's drawing on usage credits.
At 120K tokens of context, a five-minute write on Opus 5.5 costs about $0.60 and a read about $0.02. One write costs as much as 25 reads. On an API key, a six-minute coffee break turns the next $0.02 read into a $0.60 write. A one-hour write at the same size costs about $0.96, and on the API you can pay that premium to cover the gaps in your day.
Session shape and hit rate
The cache stores a prefix, so it can reuse only the part of a request that matches the previous one from the start.
A steady session appends to the end of the conversation on every turn and keeps its hit rate high. Anything that changes an earlier part of the request lowers it. Changing the tool definitions clears the whole cache, and a change to the system prompt clears it from that point on, which is almost everything.
In practice, expect a cache write when:
You pause longer than the cache lifetime;
You change effort on Amazon Bedrock, Google Cloud's Agent Platform or a gateway, where the level is part of what the cache matches;
You turn on fast mode for the first time in a conversation, which changes part of what the cache matches;
You connect or disconnect an MCP server, which can change what loads at the start of each request;
You switch models, since the new model starts from an empty cache; and
The conversation is compacted, which rewrites the history the cache matched.
So set these up when the session starts, and leave them alone while it works.
Why long sessions cost more per turn
E
首次收录 · 2026-09-26 · 10.95 分