将所有请求发送给最强的模型能获得最好的回答,但代价最高,因为对于本可由更便宜模型正确回答的请求,你仍需支付前沿模型的费率。固定的路由规则(例如按关键词或任务类型选择模型)成本较低,但随着流量变化需要重新编写。
基于置信度的升级机制介于两者之间。你要求模型对其自己的回答进行评分,然后根据该分数进行路由。得分高的回答留在便宜的模型上,得分低的回答则发送给更强的模型。本指南通过五个步骤来设置这一机制。
简而言之
基于置信度的升级机制根据模型为其自身回答报告的分数来路由每个请求,因此便宜的模型处理它确信正确的请求,而更强的模型仅处理它不确定的请求。
使用结构化输出来强制生成分数。要求包含数值型置信度字段的 JSON 模式可以从所有模型中获得相同的信号,而不是从自由文本中的模糊措辞中猜测。
使用排名而非绝对数值。0.85 并不代表校准后的 85% 正确概率。在你自己的流量数据中检查较低分数的回答是否比高分数的回答更常出错,并据此进行路由排序。
根据你的数据设定阈值。将具有代表性的流量通过便宜模型运行,查看每个分数段的错误率,并将截止点设置在错误率上升的位置。
将阈值视为需要重新审视的设置。记录分数分布、升级率和未升级回答的错误率,并在模型或流量发生变化时重新调整。
置信度分数的含义与局限性
订阅即表示你同意接收 OpenRouter 的新闻通讯:包括模型使用数据、产品更新和研究报道,每周约一封电子邮件。您可以通过每封电子邮件中的链接随时退订。请参阅我们的隐私政策。
该分数是一种自我报告。模型以生成其余回答的相同方式生成它,因此它带有相同的不确定性。两个相同的分数并不能保证具有相同的正确概率。一个提示词下的 0.85 并不等同于另一个提示词或来自其他模型的 0.85,且该数值并非校准后的概率。
一旦你验证了排名,就可以使用它。将一批你自己的请求通过便宜模型运行,对回答进行评分,并按分数分组。如果较低分数段的回答比较高分数段的回答更常出错,那么即使绝对数值不可用,这种排序也可用于路由。你寻找的不是等于 90% 正确的分数,而是区分可以发送的回答和需要二次调用的回答的排名位置。第二步就是进行这项测量。
步骤 1:使用结构化输出获取数值型置信度字段
不要从自由文本中解读不确定性。没有任何东西要求模型在不确定时使用模糊措辞,也不存在固定的模糊词汇表可供解析。
相反,让模型返回作为模式验证响应一部分的置信度字段,使用结构化输出。你传递类型为 json_schema 的 response_format,并要求同时包含 answer 和介于 0 到 1 之间的数值型 confidence:
{
"model" : "openai/gpt-5.6-luna" ,
"messages" : [
{
"role" : "user" ,
"content" : "..."
}
],
"provider" : {
"require_parameters" : true
},
"response_format" : {
"type" : "json_schema" ,
"json_schema" : {
"name" : "answer_with_confidence" ,
"strict" : true ,
"schema" : {
"type" : "object" ,
"properties" : {
"answer" : {
"type" : "string"
},
"confidence" : {
"type" : "number" ,
"description" : "回答正确的可能性,从 0(猜测)到 1(确定)。"
}
},
"required" : [ "answer" , "confidence" ],
"additionalProperties" : false
}
}
}
}
每个响应现在都携带一个数值型置信度,你的路由代码可以直接读取,无论哪个模型回答了。决定升级什么变成了数值的比较,而不是解析语言。
该模式在字段的描述中说明了0到1的范围,而不是使用最小值和最大值关键字。Anthropic的结构化输出文档将最小值和最大值等数值约束列为不支持项,因此描述形式是唯一能在各提供商之间通用的方式。strict: true 要求具有原生严格模式的提供商严格执行该模式。执行力度因提供商而异,有些提供商将模式视为强提示而非保证,因此在根据解析后的JSON进行路由之前,请务必对其进行验证。
在依赖此功能之前,请进行两项检查。首先,结构化输出支持是针对每个提供商端点设置的,而不是针对每个模型设置的,同一模型可以由支持或不支持该功能的提供商提供服务。请在模型页面中筛选出至少有一个支持端点的模型,并查看模型页面“提供商”部分中的 structured_outputs 参数。其次,在您的提供商偏好设置中设置 require_parameters: true,以便我们仅将请求路由到支持其中所有参数的端点。如果没有该标志,response_format 只是一种软性偏好。当模型具有支持的端点时,我们会将其路由至这些端点;但如果模型的某个端点均不支持它,我们仍会发送请求,而该参数将被忽略。
步骤2:根据您的错误率设定起始阈值
阈值来源于对您自身流量的测量,而非从指南中复制一个数字。使用上述模式通过您的廉价模型运行具有代表性的请求样本,记录置信度分数以及每个答案是否正确,并设置错误率开始上升时的截止点。高于该线的请求在廉价模型上解决。低于该线的请求则升级至更强的模型。
起初应采取保守策略。一开始过度升级并在稍后放宽阈值会浪费资金。而升级不足则会发送自信的错误答案。
以下是一个具体示例。假设您通过廉价模型运行了200个具有代表性的请求,并按分数段对结果进行分组。该表是内部算术一致的说明,而非目标值。您自身的分布会有所不同。
分数段 请求占比 观察到的错误率
0.95 至 1.00 41% 1%
0.85 至 0.94 27% 4%
0.70 至 0.84 18% 11%
0.50 至 0.69 9% 34%
低于 0.50 5% 61%
错误率在低于0.70时急剧上升,因此0.7是候选截止点。0.7的阈值会升级其下方的两个分数段,即14%的请求,并让其余86%的请求在廉价模型上解决。
在此样本中,廉价模型的整体错误率略低于10%。升级底部的14%将您保留的答案的错误率降低至约4%,代价是大约每七个请求中就有一个需要进行第二次调用。在自身流量上运行此分析的目的是找到您自己的截止点,而不是采用0.7。
步骤3:针对准确性、成本和延迟调整阈值
阈值是在升级量与错误率之间进行权衡。提高阈值会使更多请求获得第二次调用,从而捕获更多错误,但成本更高且速度更慢。降低阈值会使更多请求留在廉价模型上,这更快且更便宜,但会放行更多错误答案。设置阈值的位置取决于错误答案对您的产品造成的损失。有三个因素会影响这一选择。
准确性。较高的阈值会将更多边缘情况的答案发送进行第二次调用,从而减少错误答案的发布。在具体示例中,将截止点从0.7提高到0.85也会升级0.70至0.84分数段,该分数段的错误率为11%。升级比例从请求的14%上升至32%,保留答案的错误率从约4%下降至约2%。
成本。每次升级都意味着一次额外的模型调用,这发生在你已经支付的廉价调用之上。具体花费取决于你的廉价模型与升级目标模型之间的价格差距,以及你进行升级的频率。在确定目标升级率之前,请务必查看当前各模型的定价。随着新模型的发布,层级之间的差距和价格本身都会发生变化。
延迟。升级后的请求需要进行第二次往返,因此速度较慢。如果大量流量发生升级,这条慢速路径可能会挤压你的延迟预算。当你的产品有严格的响应时间上限时,这个上限可能会在成本或准确性达到极限之前,限制你能够进行的升级量。
该图表将每个候选截止值应用于示例表格。提高截止值会降低你保留的答案的错误率,但会增加支付第二次调用费用的流量比例。这三个因素并非独立变化。提高准确率阈值总是会增加升级部分的成本和延迟,因此正确的阈值应位于你对错误答案的容忍度、预算和响应时间限制三者交汇之处。
步骤 4:在代码中路由低置信度请求
一旦设定了阈值,你的代码就需要做出升级决策。它读取置信度字段,当分数低于该线时,将请求发送给更强的模型。我们不会替你做出这个决定。我们的模型回退机制仅在模型返回错误时触发,而不是在返回带有低置信度分数的有效答案时触发。有两种方式来构建这种决策逻辑。
在同一函数中重试。将调用包裹在一个函数中。向廉价模型发送请求,读取置信度,如果低于阈值,则将相同的消息发送给更强的模型并返回该答案。升级策略位于一处,因此当你更改阈值或升级目标时,只需更改一次,就不会有调用点继续做出旧的决策。
运行单独的初筛。将廉价模型的调用视为初筛。记录其答案和分数,然后仅在分数低于阈值时才调用更强的模型。这需要更多的代码,但它记录了廉价模型升级的频率,以及其分数是否仍然与实际错误相符,这是你在步骤 5 中重新调整所需的数据。
模型 ID 是字符串,因此你可以通过配置而非代码来更改任一层级。在选择廉价和强效模型时,请浏览我们的模型目录,因为随着新模型的发布,每个层级的最佳选择都会发生变化。
你可以在这两种模式之下叠加错误故障转移。传递一个 models 数组允许在第一个模型返回错误时,调用回退到列表中的下一个模型。默认情况下,任何错误都可以触发回退,包括提供商停机、速率限制、过滤模型上的审核标记以及上下文长度验证错误。我们按最终回答问题的模型对该请求进行定价,并在响应的 model 字段中返回该模型。这与你的置信度升级机制是分开的。它在发生错误时触发,而不是在有效的低置信度答案时触发,因此两者可以组合使用。回退机制确保每次调用都能存活,而你的阈值则决定何时需要更强的模型来处理一个存活的回答。
步骤 5:在生产环境中监控并重新调整
适合发布时的阈值可能会变得不再适用。从第一天起就记录三样东西。
分数分布,以便在模型或流量变化时重新运行步骤 2 的校准。
随时间变化的升级率。
未升级答案的错误率,这个数字告诉你截止值是否仍在发挥作用。
当某些情况发生变化时进行重新调整。将廉价模型或升级目标替换为新版本会改变分数分布。流量向更难或更简单的请求偏移会移动你的错误区间。成本压力可能迫使你为了更低的升级率而接受更高的错误率。阈值是一个你需要随着这些因素变化而不断调整的设定。
常见错误
将自我报告的分数视为校准后的概率。该方法适用于排序,而非字面意义上的可能性。按分数对一批答案进行排序,错误的结果会聚集在低端,但 0.9 并不意味着该答案有 90% 的概率是正确的。应从第 2 步中的“按区间划分的错误率”练习中设定阈值,而不是直接使用原始数值。
对所有任务类型使用单一阈值。回答“你们的退款政策是什么”的支持机器人和回答“如果我今天取消会被收费吗”的机器人,其答错的成本截然不同。使用单一的全局截止点会导致简单情况过度升级或高风险情况升级不足。在你知道任务类型的地方,为每种任务分配各自的阈值。
发布后不监控升级率。适合你校准批次的阈值可能会随着流量变化或底层模型更新而不再适用,且不会有任何错误提示来告诉你这一点。
结论
基于置信度的升级机制在模型报告高置信度时让请求保持在廉价模型上,仅在置信度低时才付费使用更强的模型,无需依赖关键词或任务类型规则。该循环很小。通过结构化输出强制要求数值型置信度字段。根据你自己的按分数区间的错误率寻找截止点,而不是随意选择一个数字。选择符合你追踪需求的路由模式,可以是用于策略的一个函数,也可以是第一遍处理的独立日志。然后在发布后监控未升级答案的错误率,并根据模型和流量的变化进行调整。
常见问题
什么是置信度阈值?
置信度阈值是你将请求发送给更强模型而非接受廉价模型答案的分数下限。高于该线的答案直接发出。低于该线时,请求会升级。你应根据自己的错误率数据设定这条线,而不是使用默认数值。
AI 中的置信度分数是什么?
置信度分数是一个数字,通常在 0 到 1 之间,模型在提供答案的同时报告该数字以表明其确信程度。它是自我报告,而非校准后的概率,因此 0.9 并不意味着有 90% 的正确几率。你可以测量并依赖的是排序结果。在你自己的一批请求中,检查较低分数的答案是否比高分数的答案更常出错,然后再基于分数进行路由。
设置置信度阈值的最可靠方法是什么,以便让轻量级模型处理大部分请求,仅在置信度低时 defer(委托)给高级模型?
针对你自己的流量进行测量。将具有代表性的请求样本通过轻量级模型运行,记录每个置信度分数以及答案是否正确,并按分数区间对结果进行分组。在错误率急剧上升的地方设定阈值。这样,轻量级模型将处理线以上的所有请求,只有线下低置信度的请求才会发送到高级模型。从保守设置开始,然后在你监控所保留答案的错误率时进行调整。
如何自动从小模型升级到前沿模型?
让小模型通过结构化输出返回数值型置信度字段,然后让你的代码基于该字段进行路由。当分数低于你的阈值时,将相同的请求发送给前沿模型,可以在同一函数中内联执行,或作为显式的第二次调用。OpenRouter 的模型回退是一个独立的功能。它在出现提供商停机或速率限制等错误时将调用失败并重试其他模型,而不是针对低置信度分数,因此请保持这两种机制的独立性。
参考文献
Sending every request to your strongest model gets good answers at the highest price, because you pay frontier rates for requests a cheaper model would have answered correctly. Fixed routing rules, such as picking the model by keyword or task type, cost less but need rewriting as your traffic changes.
Confidence-based escalation sits between the two. You ask the model to score its own answer, then route on that score. Answers with a high score stay on a cheap model. Answers with a low score go to a stronger one. This guide sets that up in five steps.
Tl;dr
Confidence-based escalation routes each request on a score the model reports for its own answer, so a cheap model handles the requests it is sure about and a stronger model only sees the ones it is not.
Force the score with structured outputs. A JSON schema that requires a numeric confidence field gives you the same signal from every model, instead of guessing from hedge words in free text.
Use the ranking, not the absolute number. A 0.85 is not a calibrated 85 percent chance of being correct. Check on your own traffic that lower-scored answers are wrong more often than higher-scored ones, and route on that ordering.
Set the threshold from your own data. Run representative traffic through the cheap model, look at the error rate per score band, and put the cutoff where the error rate climbs.
Treat the threshold as a setting you revisit. Log the score distribution, the escalation rate, and the error rate on non-escalated answers, and re-tune when your models or traffic change.
What the confidence score is and is not
By subscribing you agree to receive the OpenRouter newsletter: model usage data, product updates, and research reports, about one email a week. Unsubscribe anytime via the link in every email. See our Privacy Policy .
The score is a self-report. The model produces it the same way it produces the rest of its answer, so it carries the same uncertainty. Two identical scores do not guarantee the same odds of being correct. A 0.85 on one prompt is not interchangeable with a 0.85 on another prompt or from another model, and the number is not a calibrated probability.
What you can use is the ranking, once you have checked it. Run a batch of your own requests through the cheap model, grade the answers, and group them by score. If answers in the lower score bands are wrong more often than answers in the higher bands, the ordering is usable for routing even though the absolute numbers are not. You are not looking for the score that equals 90 percent correct. You are looking for the point in the ranking that separates answers you can ship from answers that need a second call. Step 2 is that measurement.
Step 1: Get a numeric confidence field with structured outputs
Do not read uncertainty out of free text. Nothing requires a model to hedge in its prose when it is unsure, and there is no fixed vocabulary of hedge words to parse.
Make the model return a confidence field as part of a schema-validated response instead, using structured outputs . You pass a response_format of type json_schema and require both an answer and a numeric confidence between 0 and 1:
{
"model" : "openai/gpt-5.6-luna" ,
"messages" : [
{
"role" : "user" ,
"content" : "..."
}
],
"provider" : {
"require_parameters" : true
},
"response_format" : {
"type" : "json_schema" ,
"json_schema" : {
"name" : "answer_with_confidence" ,
"strict" : true ,
"schema" : {
"type" : "object" ,
"properties" : {
"answer" : {
"type" : "string"
},
"confidence" : {
"type" : "number" ,
"description" : "How likely the answer is correct, from 0 (a guess) to 1 (certain)."
}
},
"required" : [ "answer" , "confidence" ],
"additionalProperties" : false
}
}
}
}
Every response now carries a numeric confidence your routing code can read directly, whichever model answered. Deciding what to escalate becomes a comparison of numbers rather than parsing language.
The schema states the 0 to 1 range in the field’s description rather than with the minimum and maximum keywords. Anthropic’s structured outputs documentation lists numerical constraints such as minimum and maximum as unsupported, so the description form is the one that works across providers. strict: true asks providers that have a native strict mode to enforce the schema exactly. Enforcement varies by provider, and some treat the schema as a strong hint rather than a guarantee, so validate the parsed JSON before you route on it.
Two checks before you rely on this. First, structured output support is set per provider endpoint, not per model, and the same model can be served by providers that do and do not support it. Filter the models page to the models with at least one supporting endpoint, and check the structured_outputs parameter in the Providers section of the model’s page. Second, set require_parameters: true in your provider preferences so we only route the request to endpoints that support every parameter in it. Without that flag, response_format is a soft preference. We route to supporting endpoints when a model has some, but if none of a model’s endpoints support it we still send the request and the parameter is ignored.
Step 2: Set a starting threshold from your own error rates
The threshold comes from measuring your own traffic, not from copying a number out of a guide. Run a representative sample of requests through your cheap model with the schema above, record the confidence score and whether each answer was correct, and set the cutoff where the error rate starts to climb. Requests above the line resolve on the cheap model. Requests below it escalate to the stronger one.
Start on the conservative side. Over-escalating at first and relaxing the threshold later costs money. Under-escalating ships confident wrong answers.
Here is a worked example. Suppose you run 200 representative requests through the cheap model and group the results by score band. The table is an illustration with internally consistent arithmetic, not a target. Your own distribution will differ.
Score band Share of requests Observed error rate
0.95 to 1.00 41% 1%
0.85 to 0.94 27% 4%
0.70 to 0.84 18% 11%
0.50 to 0.69 9% 34%
Below 0.50 5% 61%
The error rate climbs sharply below 0.70, so 0.7 is the candidate cutoff. A threshold of 0.7 escalates the two bands under it, 14 percent of requests, and lets the other 86 percent resolve on the cheap model.
In this sample the cheap model’s overall error rate is just under 10 percent. Escalating the bottom 14 percent brings the error rate on the answers you keep down to about 4 percent, at the cost of a second call on roughly one request in seven. The point of running this on your own traffic is to find your own cutoff, not to adopt 0.7.
Step 3: Tune the threshold against accuracy, cost, and latency
A threshold trades escalation volume for error rate. Raise it and more requests get a second call, which catches more errors but costs more and runs slower. Lower it and more requests stay on the cheap model, which is faster and cheaper but lets more wrong answers through. Where you set it depends on what a wrong answer costs your product. Three things shape the choice.
Accuracy. A higher threshold sends more borderline answers for a second call, so fewer errors ship. In the worked example, moving the cutoff from 0.7 to 0.85 also escalates the 0.70 to 0.84 band, which carried an 11 percent error rate. Escalation rises from 14 percent to 32 percent of requests, and the error rate on kept answers falls from about 4 percent to about 2 percent.
Cost. Every escalation is a second model call, on top of the cheap call you already paid for. What that costs depends on the price gap between your cheap model and your escalation target and on how often you escalate. Check current per-model pricing before you commit to a target escalation rate. Both the gap between tiers and the prices themselves change as new models ship.
Latency. An escalated request makes a second round trip, so it is slower. If a large share of traffic escalates, that slow path can push against a latency budget. When your product has a hard response-time ceiling, that ceiling can cap how much you escalate before cost or accuracy does.
The chart applies each candidate cutoff to the worked-example table. Raising the cutoff lowers the error rate on the answers you keep and raises the share of traffic that pays for a second call. The three factors do not move independently. Raising the threshold for accuracy always adds cost and latency on the escalated share, so the right threshold is where your tolerance for wrong answers, your budget, and your response-time limit meet.
Step 4: Route low-confidence requests in your code
Once the threshold is set, your code makes the escalation decision. It reads the confidence field and, when the score is below the line, sends the request to a stronger model. We do not make that call for you. Our model fallbacks trigger when a model returns an error, not when it returns a valid answer with a low confidence score. There are two ways to structure the decision.
Retry in the same function. Wrap the call in one function. Send the request to the cheap model, read the confidence, and if it is below the threshold, send the same messages to a stronger model and return that answer. The escalation policy lives in one place, so when you change the threshold or the escalation target you change it once and no call site keeps making the old decision.
Run a separate first pass. Treat the cheap-model call as a first pass. Log its answer and score, then call the stronger model only when the score falls below the line. This takes more code, but it records how often the cheap model escalates and whether its scores still line up with real errors, which is the data you re-tune from in Step 5.
Model IDs are strings, so you can change either tier through configuration rather than code. Browse our model catalog when you pick the cheap and strong models, since the right choice at each tier changes as new models ship.
You can layer error failover under either pattern. Passing a models array lets a call fail over to the next model in the list when the first one returns an error. By default any error can trigger the fallback, including provider downtime, rate limiting, a moderation flag on a filtered model, and a context-length validation error. We price the request at the model that ultimately answered, and return that model in the response’s model field. This is a separate mechanism from your confidence escalation. It fires on errors, not on a valid low-confidence answer, so the two compose. Fallbacks keep each call alive, and your threshold decides when a live answer needs a stronger model.
Step 5: Monitor and re-tune in production
A threshold that fits at launch can stop fitting. Log three things from day one.
The score distribution, so you can rerun the Step 2 calibration as models or traffic change.
The escalation rate over time.
The error rate on non-escalated answers, which is the number that tells you whether the cutoff is still doing its job.
Re-tune when something changes. Swapping the cheap model or the escalation target for a new version changes the score distribution. A shift in traffic toward harder or easier requests moves your error bands. Cost pressure can push you to accept a higher error rate for a lower escalation rate. The threshold is a setting you keep adjusting as those things change.
Common mistakes
Treating the self-reported score as a calibrated probability. The method works on ranking, not on literal likelihood. Sort a batch of answers by score and the wrong ones cluster toward the low end, but a 0.9 does not mean the answer is correct 90 percent of the time. Set your threshold from the error-rate-by-band exercise in Step 2, not from the raw number.
Using one threshold for every task type. A support bot answering “what is your refund policy” and one answering “will I be charged if I cancel today” do not carry the same cost of a wrong answer. A single global cutoff over-escalates the easy case or under-escalates the risky one. Where you know the task type, give each its own threshold.
Not watching the escalation rate after launch. A threshold that fit your calibration batch can stop fitting as traffic shifts or a model changes underneath you, and nothing raises an error to tell you.
Conclusion
Confidence-based escalation keeps requests on a cheap model when it reports high confidence and pays for a stronger model only when it does not, without keyword or task-type rules. The loop is small. Force a numeric confidence field with structured outputs. Find the cutoff from your own error rates by score band instead of picking a number. Pick the routing pattern that fits how you want to track it, one function for the policy or separate logs for the first pass. Then watch the error rate on non-escalated answers after launch and adjust as your models and traffic change.
Frequently asked questions
What is a confidence threshold?
A confidence threshold is the score below which you send a request to a stronger model instead of accepting the cheap model’s answer. Above the line, the answer ships as is. Below it, the request escalates. You set the line from your own error-rate data, not from a default number.
What is a confidence score in AI?
A confidence score is a number, usually between 0 and 1, that a model reports alongside its answer to signal how sure it is. It is a self-report, not a calibrated probability, so a 0.9 does not mean a 90 percent chance of being correct. What you can measure and rely on is the ordering. On a batch of your own requests, check that lower-scored answers are wrong more often than higher-scored ones before you route on the score.
What is the most reliable way to set confidence thresholds so a lightweight model handles most requests and defers to a premium model only when its confidence is low?
Measure it against your own traffic. Run a representative sample of requests through the lightweight model, record each confidence score and whether the answer was correct, and group the results by score band. Set the threshold where the error rate climbs sharply. The lightweight model then handles every request above the line, and only the low-confidence requests below it go to the premium model. Start conservative, then adjust as you watch the error rate on the answers you keep.
How do you escalate from a small model to a frontier model automatically?
Have the small model return a numeric confidence field with structured outputs, then let your code route on it. When the score is below your threshold, send the same request to the frontier model, either inline in the same function or as an explicit second call. OpenRouter’s model fallbacks are a separate feature. They fail a call over to another model on errors such as provider downtime or rate limits, not on a low confidence score, so keep the two mechanisms distinct.
References
首次收录 · 2026-10-02 · 9.95 分