“先选廉价模型”的路由策略将常规的支持问题发送给小型、低成本的模型,并将选定的复杂或不确定请求升级给更强的模型。只要低成本路径能满足您的支持需求,它就能降低为每条消息都使用高性能模型的成本。这三种应用侧的模式分别是静态规则、基于分类器的分诊以及带有置信度阈值的答案检查。
本指南对比了这三种模式,展示了 OpenRouter 处理流程中的哪些部分,通过一个可复制的请求演示了支持流程,并阐述了在添加升级机制前需要衡量的指标。
简而言之
对于狭窄且可识别的常见问题(FAQ),使用静态规则;对于以不同措辞表达的稳定类别,使用基于分类器的分诊;当希望廉价模型先尝试生成回答时,使用答案检查。
请将支持政策合规性检查保留在您的应用程序中。我们的模型回退机制处理请求错误。成功但错误的答案需要在您的代码中进行单独的升级决策。
在答案接受率、包括被丢弃尝试在内的总成本、升级率以及完整对话延迟方面,将“先选廉价模型”与“仅选廉价模型”和“仅选更强模型”进行对比。
仅在评估显示升级机制能修复失败的地方添加升级功能。如果一次升级替换了原本已可接受的廉价答案,只会增加成本和延迟,而不会改善响应质量。
当更便宜的模型适用于常规常见问题时
使用一个高性能模型可以保持初始集成的简单性。如果常规常见问题在更便宜的模型上能满足您的质量要求,那么这些流量就可以成为降低模型支出的候选对象。使用我们模型目录中的候选模型和当前定价,针对同一份经批准的文档测试这两个模型。
分类器调用和第二次尝试也会消耗时间和预算。根据您在生成前能识别的内容以及生成后能检查的内容来选择模式。
三种路由模式
Anthropic 关于构建有效代理的指南将路由描述为对输入进行分类并将其引导至专门的后续任务,简单问题发送给较小的模型,复杂问题发送给能力更强的模型。以下构建和维护估算值是我们针对小型支持应用的规划判断。
| 方法 | 决策点 | 构建工作量 | 维护负担 | 适用的起始工作负载 |
|---|---|---|---|---|
| 静态规则 | 在生成前匹配意图或短语 | 低 | 随措辞和政策变化更新规则 | 狭窄、可识别的常见问题类别 |
| 基于分类器的分诊 | 轻量级分类器选择模型或工作流 | 中 | 维护标注示例并检查误路由 | 以不同语言表达的稳定类别 |
| 带有置信度阈值的答案检查 | 接受廉价答案或请求再次尝试 | 中高 | 维护政策检查,并根据正确性验证置信度分数 | 具有明确、可测试要求的答案 |
这些模式可以共存。规则或分类决定请求的起点,答案检查决定其输出是否可接受。在您根据自有问题的正确性对其进行检查之前,模型自报告的置信度并非正确性的概率,因此在获得该数据之前,请将其视为排名信号。
您的应用程序拥有哪些功能以及 OpenRouter 处理哪些部分
拥有基于质量的路由意味着维护规则或分类器、接受检查、升级决策和监控。这些可以存在于您的支持应用内部,无需单独的路由服务。
如果您希望我们选择初始模型,我们的 Auto Router(通过 "model": "openrouter/auto" 选择)会根据任务类型对每个提示进行分类,并从 OpenRouter 用户在该任务类型上花费最多的模型中选择一个模型。有序模型回退列表(作为 models 数组发送)会在出现错误(如速率限制、提供商停机或内容审核拒绝)时重试下一个模型。支持政策合规性检查仍由您负责。
如果您希望我们选择初始模型,请选择自动路由(Auto Router)。如果您清楚偏好的模型并需要错误恢复机制,请选择有序的备用模型列表。对于“先低成本、后质量提升”的策略,您的应用程序会检查候选答案,并在检查失败时请求另一个回答。
以下流程展示了这种职责划分。
OpenRouter 上的一个有效支持流程示例
步骤 1:在选择模型之前定义预期结果
HarborDesk 是一个虚构的 SaaS 产品,拥有固定的常见问题解答(FAQ)。其机器人可以解释政策并推荐支持服务,但无法检查账户或进行修改。可接受的回答应能回答问题、请求澄清或推荐人工支持。
对于您能够可靠识别且具有固定答案的意图,请使用已批准的模板。单独测试意图选择,以确保正确的措辞对应正确的问题。需要解读的请求将进入下方的生成答案路径。
步骤 2:在升级之前检查生成的答案
对于需要生成的请求,低成本模型提供第一个候选答案和分类信号。您的应用程序拥有接受决策权。
考虑一位询问如何取消订阅并申请退款的客户。根据 HarborDesk 的政策,所有者应在“设置 > 账单”中取消订阅,而账单团队负责审核退款请求。
接受一个解释了这两个步骤但未承诺退款的低成本答案。
拒绝一个将所有内容发送至账单部门且遗漏所有者取消操作的答案。如果已有涵盖该请求的已批准模板,请使用它;或者,如果您的评估显示更强的模型能修正此遗漏,则请求使用更强模型再次尝试。
如果客户要求机器人批准退款,请推荐账单支持。更强的模型无法提供此类权限。
检查政策内容和预期操作。仅凭有效的 JSON 并不能确立正确性。保留对话上下文,对更强模型的答案执行相同的检查,并在一次升级后停止。
步骤 3:配置错误备用并记录请求成本
此代码片段发送一个带有有序备用模型列表的请求。在运行之前安装 requests 库并设置 OPENROUTER_API_KEY 环境变量。示例中的简短虚构 FAQ 提供了上下文。
import os
import requests
messages = [
{
"role" : "system" ,
"content" : (
"Answer only from this fictional HarborDesk FAQ. "
"Owners can cancel under Settings > Billing. "
"Billing support reviews refund requests. "
"You cannot inspect accounts, cancel plans, or approve refunds. "
"Ask for clarification when necessary."
),
},
{ "role" : "user" , "content" : "How do I cancel my subscription and request a refund?" },
]
response = requests.post(
"https://openrouter.ai/api/v1/chat/completions" ,
headers = { "Authorization" : "Bearer " + os.environ[ "OPENROUTER_API_KEY" ]},
json = {
"models" : [ "openai/gpt-4.1-mini" , "openai/gpt-4.1" ],
"messages" : messages,
"max_tokens" : 512 ,
"stream" : False ,
},
timeout = 60 ,
)
response.raise_for_status()
data = response.json()
if "error" in data:
raise RuntimeError ( f "Completion failed: { data[ 'error' ] } " )
choices = data.get( "choices" , [])
if not choices:
raise RuntimeError ( "Completion returned no choices" )
choice = choices[ 0 ]
if choice.get( "error" ) is not None or choice.get( "finish_reason" ) == "error" :
raise RuntimeError ( f "Completion failed: { choice.get( 'error' , 'unknown error' ) } " )
print ({
"id" : data[ "id" ],
"model" : data[ "model" ],
"answer" : choice[ "message" ][ "content" ],
"cost" : data.get( "usage" , {}).get( "cost" ),
})
models 数组会在遇到符合条件的错误(如速率限制或提供商停机)时尝试下一个模型。成功但不正确的答案不会触发此机制。质量提升需要应用程序检查以及单独的更强模型请求。
这两个错误检查的存在是因为当模型开始处理请求但在生成输出时失败,我们可以返回 HTTP 200。错误对象出现在响应体的顶层,或者在 choices[0] 内部,且 finish_reason 设置为 error(当存在部分输出时)。错误和调试参考文档描述了这两种结构。应将两者都视为失败,而不是将部分输出作为答案打印出来。
该代码片段会打印一个未经验证的候选结果以供检查。usage.cost 字段是我们为此次请求收取的费用(以积分计),model 是生成响应的模型,这在触发回退机制时很重要。请对每次请求记录这两项数据。当你的应用程序进行升级时,请汇总两次调用的成本,包括被丢弃的廉价草稿模型的费用。如果响应中没有包含 cost 值,请在与活动页面核对之前,将成本记录为未知(unknown)而不是零。usage 会计食谱描述了 usage 对象。
衡量升级率和每张已解决工单的成本
对于始终优先尝试廉价模型的流水线,请使用以下公式计算每次请求的预期模型成本:
expected_cost = cost_cheap + cost_check_1
+ escalation_fraction * (cost_strong + cost_check_2)
在 check 项中包含任何评估器或分类器调用的成本,并在升级的请求中使用更强模型的成本。将结果与仅使用廉价模型和仅使用更强模型的情况进行比较,同时保持相同的质量目标和延迟限制。
针对每种路由策略跟踪这些指标。
升级率。采用更强模型尝试的请求所占的比例。
被接受的错误答案。通过了你的检查但在人工审核中失败的廉价答案。
被升级的正确答案。已经符合你标准的答案却被意外升级的情况。每一次这样的案例都会增加一次更强模型的计费以及第二次调用的延迟,而并未改善响应结果。
交接的适当性。当请求需要其不具备权限的人工支持时,机器人是否推荐了人工支持。
完整轮次延迟。从客户发送消息到返回答案的时间,包括在升级请求中的两个模型阶段。
高升级率值得检查被路由的问题。但这本身并不能证明阈值设置错误。
对于每张已解决工单的成本,请将一个工单群体(包括未解决的工单)的所有模型支出,除以在规定的重新打开窗口期内符合既定解决定义的工单数量。在测试集上对答案接受度的衡量反映的是响应质量,而非客户问题的实际解决情况,因此请保持这两个指标分开。
在生产环境启用升级功能之前,请在保留的支持问题上运行此比较。包括简单的常见问题解答(FAQ)、模糊的请求、故障排除以及需要人工交接的请求,并根据预定义的内容标准和预期操作对每个答案进行评分。只有在将失败的答案转化为成功答案,且其成本和延迟在可接受范围内时,升级才值得存在。
结论
从批准的模板和针对狭窄 FAQ 的廉价模型基线开始。根据你的工作负载选择规则、分类或答案检查,然后仅在相同的质量目标和延迟限制下与更强模型进行比较。仅在能够修复足够多的失败以证明成本和延迟合理的情况下添加升级功能。当支持政策发生变化时,请重新审视你的检查机制。
常见问题
我可以避免构建自定义路由服务吗?
你可以将路由保留在你的支持应用程序内部。如果你希望我们选择初始模型,请使用 Auto Router;如果希望在请求出错后重试另一个模型,请使用 models fallback 数组。特定于支持的接受性检查(例如答案是否符合你的取消政策)仍然属于你的应用程序。请将模型 ID 保存在配置中,并保持提示词和评估数据的可移植性,以便将来可以更换模型。
廉价模型能否在生成过程中咨询更强模型?
不要通过本指南中的流程。廉价模型会率先完成候选答案,随后你的应用程序决定是否请求更强模型的回复。在对话轮次中咨询另一个模型需要在代码中显式地调用该更强模型的工具或编排步骤,并将其输出返回给正在进行的流程。模型的后备数组并不会创建这种机制,它仅在出错时尝试在其他模型上重试请求。
我应该选择哪个 FAQ 分流模型?
从满足你保留支持问题质量目标的最低成本候选者开始,包括模糊的请求和政策例外情况。测量完整对话轮次的延迟,而不是单次调用的延迟,因为最快的模型调用仍可能因经常需要第二次尝试而导致较慢的支持响应。从一组重复出现的问题开始,并随着评估结果展示哪些路径有效而逐步扩展。
参考文献
Cheap-first model routing sends routine support questions to a small, inexpensive model and escalates selected hard or uncertain requests to a stronger model. It can reduce the cost of using a capable model for every message, as long as the cheaper path meets your support requirements. The three application-side patterns are static rules, classifier-based triage, and answer checks with confidence thresholds.
This guide compares the three patterns, shows which parts of the flow OpenRouter handles, walks through a support flow with a copyable request, and sets out what to measure before you add escalation.
Tl;dr
Use static rules for narrow, recognizable FAQs, classifier-based triage for stable categories expressed in varied wording, and answer checks when a cheap model should attempt the response first.
Keep support-policy acceptance checks in your application. Our model fallbacks handle request errors. A successful but incorrect answer needs a separate escalation decision in your code.
Compare cheap-first against cheap-only and stronger-only on answer acceptance, total cost including discarded attempts, escalation rate, and full-turn latency.
Add escalation only where your evaluation shows it repairs failures. An escalation that replaces an already acceptable cheap answer adds cost and latency without improving the response.
When a cheaper model fits routine FAQs
One capable model keeps the initial integration simple. If routine FAQs meet your quality requirements on a cheaper model, that traffic becomes a candidate for lower model spend. Test both models against the same approved documentation, using candidates from our model catalog and current pricing .
Classifier calls and second attempts also consume time and budget. Choose a pattern based on what you can identify before generation and what you can check afterward.
Three routing patterns
Anthropic’s guidance on building effective agents describes routing as classifying an input and directing it to a specialized follow-up task, with easy questions sent to smaller models and hard questions to more capable ones. The build and maintenance estimates below are our planning judgments for a small support application.
Approach Decision point Build effort Maintenance burden Useful starting workload
Static rules Match intent or phrases before generation Low Update rules as wording and policies change Narrow, recognizable FAQ categories
Classifier-based triage A lightweight classifier selects a model or workflow Medium Maintain labeled examples and inspect misroutes Stable categories expressed in varied language
Answer checks with confidence thresholds Accept the cheap answer or request another attempt Medium to high Maintain policy checks and validate confidence scores against correctness Answers with explicit, testable requirements
These patterns can coexist. Rules or classification choose where a request starts. Answer checks decide whether its output is acceptable. A model’s self-reported confidence is not a probability of correctness until you have checked it against correctness on your own questions, so treat it as a ranking signal until you have that data.
What your application owns and what OpenRouter handles
Owning quality-based routing means maintaining rules or a classifier, acceptance checks, escalation decisions, and monitoring. These can live inside your support application without a separate routing service.
If you want us to select the initial model, our Auto Router , selected with "model": "openrouter/auto" , classifies each prompt by task type and picks a model from the ones OpenRouter users spend on for that task type. An ordered model fallback list, sent as the models array, retries the request on the next model after an error such as a rate limit, provider downtime, or a content moderation refusal. Support-policy acceptance checks remain your responsibility.
Choose the Auto Router if you want us to select the initial model. Choose an ordered fallback list if you know your preferred models and need error recovery. For cheap-first quality escalation, your application checks the candidate and requests another answer when the check fails.
The following flow shows that division of responsibility.
A worked support flow on OpenRouter
Step 1: Define the outcome before choosing a model
HarborDesk is a fictional SaaS product with a fixed FAQ. Its bot can explain policies and recommend support, but it cannot inspect accounts or make changes. An acceptable turn answers the question, asks for clarification, or recommends human support.
Use approved templates for intents you recognize reliably and that have fixed answers. Test intent selection separately so the correct wording reaches the correct question. Requests that need interpretation enter the generated-answer path below.
Step 2: Check generated answers before escalating
For requests that need generation, the cheap model provides the first candidate and the triage signals. Your application owns the acceptance decision.
Consider a customer who asks how to cancel a subscription and request a refund. In HarborDesk’s policy, owners cancel under Settings > Billing, and billing staff review refund requests.
Accept a cheap answer that explains both steps without promising a refund.
Withhold an answer that sends everything to billing and omits owner cancellation. Use an approved template if one covers the request, or request a stronger-model attempt if your evaluation shows the stronger model fixes this omission.
Recommend billing support if the customer asks the bot to approve the refund. A stronger model cannot supply that authority.
Check the policy content and the intended action. Valid JSON alone does not establish correctness. Preserve conversation context, run the same checks on the stronger answer, and stop after one escalation.
Step 3: Configure error fallback and record request cost
This snippet sends one request with an ordered fallback list. Install requests and set OPENROUTER_API_KEY before running it. The short fictional FAQ supplies the context for the example.
import os
import requests
messages = [
{
"role" : "system" ,
"content" : (
"Answer only from this fictional HarborDesk FAQ. "
"Owners can cancel under Settings > Billing. "
"Billing support reviews refund requests. "
"You cannot inspect accounts, cancel plans, or approve refunds. "
"Ask for clarification when necessary."
),
},
{ "role" : "user" , "content" : "How do I cancel my subscription and request a refund?" },
]
response = requests.post(
"https://openrouter.ai/api/v1/chat/completions" ,
headers = { "Authorization" : "Bearer " + os.environ[ "OPENROUTER_API_KEY" ]},
json = {
"models" : [ "openai/gpt-4.1-mini" , "openai/gpt-4.1" ],
"messages" : messages,
"max_tokens" : 512 ,
"stream" : False ,
},
timeout = 60 ,
)
response.raise_for_status()
data = response.json()
if "error" in data:
raise RuntimeError ( f "Completion failed: { data[ 'error' ] } " )
choices = data.get( "choices" , [])
if not choices:
raise RuntimeError ( "Completion returned no choices" )
choice = choices[ 0 ]
if choice.get( "error" ) is not None or choice.get( "finish_reason" ) == "error" :
raise RuntimeError ( f "Completion failed: { choice.get( 'error' , 'unknown error' ) } " )
print ({
"id" : data[ "id" ],
"model" : data[ "model" ],
"answer" : choice[ "message" ][ "content" ],
"cost" : data.get( "usage" , {}).get( "cost" ),
})
The models array tries the next model after an eligible error, such as a rate limit or provider downtime. A successful but incorrect answer does not trigger it. Quality escalation requires an application check and a separate stronger-model request.
The two error checks exist because we can return HTTP 200 when the model started processing the request and failed while producing output. The error object appears at the top level of the body, or inside choices[0] with finish_reason set to error when partial content exists. The errors and debugging reference describes both shapes. Treat both as failures rather than printing partial output as an answer.
The snippet prints an unvalidated candidate for inspection. The usage.cost field is the amount we charged for the request in credits, and model is the model that produced the response, which matters when the fallback fired. Log both for every request. When your application escalates, sum both calls, including the discarded cheap draft. If a response arrives without a cost value, record the cost as unknown rather than zero until you reconcile it against your activity page. The usage accounting cookbook describes the usage object.
Measure escalation rate and cost per resolved ticket
For a pipeline that always tries the cheap model first, use this formula for the expected model cost per request.
expected_cost = cost_cheap + cost_check_1
+ escalation_fraction * (cost_strong + cost_check_2)
Include the cost of any evaluator or classifier calls in the check terms, and use the stronger model’s cost on escalated requests. Compare the result against cheap-only and stronger-only, including their checks, at the same quality target and latency limit.
Track these measures for each routing policy.
Escalation rate. The share of requests that took a stronger-model attempt.
Incorrect answers accepted. Cheap answers that passed your checks but failed human review.
Correct answers escalated. Cheap answers that already met your criteria but were escalated anyway. Each one adds a stronger-model charge and a second call’s latency without improving the response.
Handoff appropriateness. Whether the bot recommended human support when the request needed authority it does not have.
Full-turn latency. Time from the customer’s message to the returned answer, including both model stages on escalated requests.
A high escalation rate warrants inspecting the routed questions. It does not by itself prove the threshold is wrong.
For cost per resolved ticket, divide all model spending for a ticket cohort, including unresolved tickets, by the number of tickets that met a stated resolution definition within a stated reopen window. Answer acceptance on a test set measures response quality, not customer resolution, so keep the two metrics separate.
Run this comparison on held-out support questions before you enable escalation in production. Include straightforward FAQs, ambiguous requests, troubleshooting, and requests that need a human handoff, and grade each answer against predefined content criteria and the expected action. Escalation earns its place only where it converts failing answers into passing ones at a cost and latency you accept.
Conclusion
Begin with approved templates and a cheap-model baseline for a narrow FAQ. Choose rules, classification, or answer checks based on your workload, then compare against stronger models only at the same quality target and latency limit. Add escalation only where it fixes enough failures to justify the cost and delay. Revisit your checks when support policies change.
Frequently asked questions
Can I avoid building a custom routing service?
You can keep routing inside your support application. Use the Auto Router if you want us to pick the initial model, and use the models fallback array so we retry another model after a request error. The support-specific acceptance checks, such as whether an answer follows your cancellation policy, still belong in your application. Keep model IDs in configuration and keep your prompts and evaluation data portable so you can change models later.
Can a cheap model consult a stronger model mid-generation?
Not through the flow in this guide. The cheap model finishes a candidate answer first, then your application decides whether to request a stronger-model answer. Consulting another model during a turn requires an explicit tool or orchestration step in your code that calls the stronger model and returns its output to the ongoing workflow. The models fallback array does not create that mechanism. It only retries the request on another model after an error.
Which FAQ triage model should I choose?
Start with the cheapest candidate that meets your quality target on held-out support questions, including ambiguous requests and policy exceptions. Measure full-turn latency rather than single-call latency, because the fastest model call can still produce a slower support response if it often needs a second attempt. Begin with a small set of recurring questions and expand as your evaluation shows which routes work.
References
| 刊期 | 得分 | 排名 | 结果 |
|---|---|---|---|
| 2026-10-05 | 8.77 | 17 | 入选 |
| 2026-10-04 | 9.22 | 24 | 未入选 |
| 2026-10-03 | 9.95 | 30 | 未入选 |