通过排行榜排名来选择代理模型,意味着对于本可由更便宜模型以相同准确率完成的任务,你却支付了前沿模型的高昂价格。需要回答的问题不是哪个模型的得分最高,而是哪个模型是足以胜任眼前任务的、成本最低的那个。
本指南提供了一个三步框架来做出这一决策。你设定任务所需的质量门槛,在自有样本上测量每个质量点的成本,然后选择那个以较大余量通过门槛的最低成本模型。
简而言之
在比较模型之前,为每项任务设定质量门槛。无论多便宜,低于门槛的模型将被直接淘汰。
在你的 20 到 50 个自有样本上运行一个低成本、一个中端和一个前沿模型,使用同一套评分标准进行打分,然后用成本除以得分得出每个质量点的成本。
选择那个以超过你观测到的运行间分数波动余量通过门槛的最低成本模型。
从每条响应中的 usage.cost 字段读取成本,而不是将列出的费率乘以预估的令牌数量。
当候选模型或其价格发生变化时,重新进行比较。
排行榜排名无法告诉你的事情
通过订阅,你同意接收 OpenRouter 的新闻通讯:包括模型使用数据、产品更新和研究报告,大约每周一封电子邮件。您可以通过每封电子邮件中的链接随时取消订阅。请参阅我们的隐私政策。
排行榜将模型在各种与你无关的任务上的结果进行了平均。在编码任务中排名第一的模型,在你的结构化数据提取任务中未必也是第一;而在推理基准测试上表现平平的中端模型,对于你的常见问题解答流量而言,其准确率可能足以满足你的要求。
代理任务通常很狭窄,例如分类工单、提取一个字段,或在模型不确定时进行升级处理。一个更便宜的模型在狭窄任务上能否达到与前沿模型相同的准确率,这是一个需要测量的问题,而排行榜并不会为你完成这一测量。
代理的成本也不止是一次提示和一次响应。单次聊天完成仅计费一次。代理则为每次工具调用、每个中间步骤以及每次重试付费。在一个三步循环中,它在返回答案之前至少支付了三次每令牌价格。如果你按排行榜排名来选择,你可能会为原本可由更便宜模型以相同准确率完成的任务,支付三倍的前沿模型价格。
第一步:定义任务所需的质量门槛
在比较任何模型之前,先决定什么对于此任务来说算是“足够好”。这个门槛因工作而异,它是后续每个步骤所依据的过滤器。
如果错误的答案会带来责任风险,例如在合规、医疗或法律审查中,请设定较高的门槛并接受每次请求更高的成本。如果任务是大规模支持或聊天,整体结果比任何单次响应更重要。一个能正确解决 90% 常规请求并干净地升级剩余 10% 的更便宜模型,对于该任务来说可能是可以接受的。你决定门槛的位置,框架则据此进行测量。
对延迟敏感的工作,例如欺诈检查和实时聊天,增加了第三个约束条件。一个既便宜又准确但速度太慢而无法胜任任务的模型,在考虑成本或准确率之前就会被淘汰。
速度、成本和准确率无法同时最大化。确定你的任务有哪些约束条件,候选名单就会缩小。
如果你不知道某项任务处于什么位置,一个起点是对所有任务运行中端模型,然后根据测量结果显示问题的地方对任务进行分类。中端模型表现超出需求的那些任务可以降级到更便宜的模型;未能达到门槛的任务则升级。由测量决定每项任务所需的层级,而不是靠事先的猜测。
第二步:在自有样本上测量每个质量点的成本
提取你的代理实际会看到的文档、工单和提示。使用你自己的流量数据,而不是公开数据集或基准测试示例。目的是测量你自身的任务。
在这些示例上运行一个低价选项、一个中端选项和一个前沿选项。使用一套统一的标准对它们进行评分。对于分类等确定性任务,使用与固定答案的精确匹配,这适用于只有一个正确答案的场景。对于开放式任务,使用大语言模型作为裁判来评估质量。我们的“大语言模型即裁判”指南详细介绍了如何设置这一机制。当每个候选者返回相同形状的输出时,评分会更容易,这正是结构化输出和 response_format 的作用所在。
将模型每 1,000 次请求的成本除以其质量得分,即可得到每个质量点的成本,从而可以直接跨候选者和不同规模的测试集进行比较。运行一个简短的脚本,该脚本交换一个模型字符串,并从每个响应的 usage 对象中读取每次请求的成本。如使用量会计文档所述,我们在每个非流式响应和每个流式响应的最后一条消息中都返回 usage 对象,无需任何额外的请求参数。usage.cost 是我们向您的账户收取的该次请求的费用(以美元计)。OpenAI Python SDK 保留了其未定义的响应字段,因此 r.usage.cost 可作为属性使用。
import json
import os
from openai import OpenAI
client = OpenAI(
base_url = "https://openrouter.ai/api/v1" ,
api_key = os.environ[ "OPENROUTER_API_KEY" ],
)
ROUTING_PROMPT = open ( "routing_prompt.txt" ).read()
test_set = json.load( open ( "test_set.json" )) # 20-50 of your own {"ticket": ..., "label": ...} examples
candidates = [
"openai/gpt-5.6-luna" ,
"google/gemini-3.7-flash" ,
"anthropic/claude-opus-5" ,
]
for model in candidates:
total_cost, correct = 0.0 , 0
for ex in test_set:
r = client.chat.completions.create(
model = model,
messages = [
{ "role" : "system" , "content" : ROUTING_PROMPT },
{ "role" : "user" , "content" : ex[ "ticket" ]},
],
)
total_cost += r.usage.cost
answer = (r.choices[ 0 ].message.content or "" ).strip()
correct += int (answer == ex[ "label" ])
score = correct / len (test_set)
points = score * 100
cost_per_1k = total_cost / len (test_set) * 1000
cost_per_point = cost_per_1k / points if points else float ( "inf" )
print ( f " { model } : score= { score :.0% } cost/1K=$ { cost_per_1k :.2f } cost/point=$ { cost_per_point :.4f } " )
响应可能不包含任何文本内容,例如当模型拒绝回答时,因此脚本将缺失的内容视为空答案,将其计为错误,并将其成本计入总额。得分为零的模型没有每个质量点的成本,因此脚本会打印 inf 而不是除以零。
以下是支持路由任务的一个工作示例。假设每个请求发送约 700 个输入令牌(包括工单和简短的系统提示),并返回约 150 个输出令牌。每 1,000 次请求的成本为 (700 / 1M × 输入价格) + (150 / 1M × 输出价格),再乘以 1,000。每个质量点的成本是该成本除以质量得分。
价格为 2026-09-18 我们模型目录中 GPT-5.6 Luna、Gemini 3.7 Flash 和 Claude Opus 5 的列出的费率。这些是目录列出的模型级别价格。各个提供商端点以及 flex 和 priority 服务层级按各自费率计费,且对于提示令牌数达到或超过 272,000 的请求,GPT-5.6 Luna 收取更高的费率。令牌数量和质量得分仅为示例说明。您的实际运行将替换这些数值。
模型 输入/输出价格(每百万令牌) 每 1,000 次请求成本 质量得分 每个质量点成本 是否达到 85% 门槛?
openai/gpt-5.6-luna $0.20 / $1.20 $0.32 82 $0.0039 否(−3)
google/gemini-3.7-flash $0.75 / $3.75 $1.09 89 $0.0122 是(+4)
anthropic/claude-opus-5 $5.00 / $25.00 $7.25 97 $0.0747 是(+12)
按照框架运行的顺序阅读表格。首先检查门槛。在 85% 的阈值下,GPT-5.6 Luna 得分为 82,因此被淘汰,其低每个质量点成本已无关紧要。低于门槛的模型无论每个质量点多么便宜都会被取消资格。剩下的候选者是 Gemini 3.7 Flash 和 Claude Opus 5。
在这两者之间,选择更便宜的那个。Gemini 3.7 Flash 每千次请求收费 1.09 美元,而 Claude Opus 5 收费 7.25 美元,尽管后者完成的任务并不要求如此高的分数。因此,Gemini 3.7 Flash 是更好的选择,它以 4 分的优势胜出。
情况改变,答案也会随之改变。如果路由错误意味着错过服务等级协议(SLA),且你将达标线设定为 95%,那么只有 Claude Opus 5 能达标,你需要支付 7.25 美元。目标不是选择得分最高的模型,而是找出哪个模型能以最低成本达到你的要求,并观察当你超越该要求时多支付了多大的差距成本。
第三步:选择以余量达标的最低成本模型
选择以余量达标的最低成本模型,而不是单纯获胜的模型。余量很重要,因为数据是动态变化的。
随着提供商更新权重、你的流量随时间推移发生变化以及新版本改变整体格局,模型得分会发生漂移。应从观察中设定余量,而不是选择一个固定的百分比。对候选模型进行多次运行,或在新的流量切片上运行,记录每次运行之间分数的波动幅度,并要求获胜者达标的幅度超过这一观察到的波动范围。
如果一个模型仅在单个小样本中勉强达到达标线,那么在你就此做出承诺之前,它必须达到更高的内部目标,以便普通的波动不会导致其在生产环境中低于达标线。
应用于几个任务时,该框架如下所示。成本数据使用与上述表格相同的列出价格,并对应每行中标注的输入和输出 token 大小。质量达标线仅为示例,你记录的成本来自 usage.cost。
| 任务 | 质量达标线 | 达标的最低层级 | 每千次请求成本 |
|---|---|---|---|
| 支持分类(700 输入/150 输出 token) | 80% | 廉价型, openai/gpt-5.6-luna | $0.32 |
| 代码审查(4,000 输入/1,000 输出 token) | 88% | 中端型, google/gemini-3.7-flash | $6.75 |
| 合规性审查(3,000 输入/600 输出 token) | 95% | 前沿型, anthropic/claude-opus-5 | $30.00 |
每一行的方法相同,但答案不同是因为达标线不同。这种衡量的权衡关系在你测量当天是真实的。当发布新模型时请重新检查,因为昨天达标的模型现在可能与仅略低于达标线的模型成本相同。
当模型或价格发生变化时,请重新比较
新模型经常发布,现有模型的价格也会变化。在 2026 年 5 月至 9 月期间,我们的目录中增加了四个版本的 Gemini Flash:2026-05-19 的 Gemini 3.5 Flash、2026-07-21 的 Gemini 3.6 Flash、2026-08-13 的 Gemini 3.7 Flash 以及 2026-09-02 的 Gemini 3.8 Flash。几个月前进行的比较可能已经过时,因此每当候选模型更新或其定价发生变化时,都应重新运行比较。
在我们这边,重新运行的成本很低。我们提供一个兼容 OpenAI 的 API,因此切换模型只需更改配置。基础 URL、密钥和 SDK 保持不变。你只需将模型字符串从 openai/gpt-5.6-luna 更改为 anthropic/claude-opus-5,上述脚本即可针对新模型运行。
你无需为每个想要测试的提供商编写新的集成,这使得在发布当周对新版本重新运行示例变得切实可行。正是这行代码的替换使得多智能体设置中的模型路由成为可能。
两件事使该框架扎根于真实数据。我们在你做出承诺之前在模型目录中展示各模型的定价,因此你比较的成本部分始于实时数据。而且,由于每个响应都报告 usage.cost,你测量的支出是我们实际计费的费用,而非来自费率表的估算值。
常见问题
AI 模型的成本与质量权衡是什么?
这是指请求成本与模型在你的任务上表现好坏之间的权衡。速度是第三个约束条件。一个便宜且准确但速度太慢无法满足任务需求的模型仍然不可用。选择模型意味着决定任务需要多少质量并为此付费,而不是为可用的最高得分付费。
AI 代理的最佳模型是什么?
不存在唯一的最佳模型。适合智能体的正确模型取决于其任务必须达到的质量门槛、必须响应的延迟以及需要处理的数据量。一个能满足合规审查门槛的模型,其成本高于支持分诊等任务所需;而足够便宜以用于分诊的模型,可能无法达到合规审查的质量要求。
AI 模型的智能体工作负载如何定价?
我们目录中的文本模型列出了每个 token 的提示费率和补全费率。部分端点还列出了缓存输入、推理 token、每次请求费用以及灵活或优先服务层级的费率。智能体工作负载在首次请求的基础上增加了工具调用和重试,因此单次补全的 token 价格会低估单个任务的真实成本。请在完整的智能体运行周期内衡量成本,并从每个响应中的 usage.cost 字段读取数据。
多智能体设计比单智能体更贵还是更便宜?
这取决于工作如何拆分。多智能体设计可以将路由和分类任务发送给廉价模型,仅对需要更高能力的请求调用更昂贵的模型,这比将所有请求都发送给昂贵模型的成本更低。但这也增加了请求数量,因此应衡量完整的运行过程,而不是假设拆分一定能省钱。
结论
明确质量门槛,使用自己的示例测量完整的运行结果,并选择成本最低且能超越你在不同运行间观察到的波动范围的质量门槛的模型。保持脚本和示例集不变,以便在模型、价格或你的流量发生变化时重新运行比较。
参考资料
Picking a model for an agent by leaderboard rank pays frontier prices for tasks that a cheaper model may handle at the same accuracy. The question to answer is not which model scores highest, but which model is the cheapest one that is good enough for the task in front of you.
This guide is a three-step framework for making that call. You set the quality bar the task needs, measure cost per quality point on your own examples, and pick the cheapest model that clears the bar with margin.
Tl;dr
Set a quality bar per task before you compare models. A model below the bar is disqualified no matter how cheap it is.
Run a cheap, a mid-tier, and a frontier model on 20 to 50 of your own examples, score them with one rubric, and divide cost by score to get cost per quality point.
Pick the cheapest model that clears the bar by more than the score swing you observe between runs.
Read cost from the usage.cost field in each response rather than multiplying listed rates by estimated token counts.
Rerun the comparison when a candidate model or its price changes.
What a leaderboard rank does not tell you
By subscribing you agree to receive the OpenRouter newsletter: model usage data, product updates, and research reports, about one email a week. Unsubscribe anytime via the link in every email. See our Privacy Policy .
A leaderboard averages a model’s results across tasks that have nothing to do with yours. The model that ranks first for coding is not necessarily first at your structured-data extraction task, and a mid-tier model that looks unremarkable on a reasoning benchmark may answer your FAQ traffic accurately enough for your bar.
Agent tasks are often narrow, such as classifying a ticket, extracting one field, or escalating when the model is unsure. Whether a cheaper model reaches the same accuracy as a frontier model on a narrow task is a measurement, and the leaderboard does not make it for you.
Cost for an agent is also more than one prompt and one response. A single chat completion is billed once. An agent pays for every tool call, every intermediate step, and every retry. A three-step loop pays the per-token price at least three times before it returns an answer. If you pick by leaderboard rank, you can pay frontier prices three times over for a task that a cheaper model would have handled at the same accuracy.
Step 1: Define the quality bar the task needs
Before you compare anything, decide what good enough means for this task. The bar is different for every job, and it is the filter every later step runs through.
If a wrong answer is a liability, as in compliance, healthcare, or legal review, set the bar high and accept the higher cost per request. If the task is high-volume support or chat, the aggregate outcome matters more than any single response. A cheaper model that resolves 90% of routine requests correctly and escalates the remaining 10% cleanly may be acceptable for that task. You decide where the bar sits, and the framework measures against it.
Latency-sensitive work, such as fraud checks and live chat, adds a third constraint. A model that is cheap and accurate but too slow for the task is disqualified before cost or accuracy come up.
Speed, cost, and accuracy can’t all be maximized at once. Decide which constraints your task has, and the candidate list narrows.
If you don’t know where a task sits, one starting point is to run a mid-tier model for everything, then split tasks by where the measurements show a problem. Tasks where the mid-tier model is more than you need move to something cheaper. Tasks where it misses the bar move up. Measurement decides which tier each task needs instead of a guess made up front.
Step 2: Measure cost per quality point on your own examples
Pull the documents, tickets, and prompts your agent will really see. Use your own traffic, not a public dataset and not benchmark examples. The point is to measure your task.
Run a cheap option, a mid-tier option, and a frontier option across those examples. Score them with one consistent rubric. For deterministic tasks like classification, use exact match against a fixture, which fits when exactly one value is correct. For open-ended tasks, use an LLM as a judge to score quality. Our LLM-as-a-judge guide walks through setting that up. Scoring is easier when every candidate returns the same shape of output, which is what structured outputs and response_format are for.
Divide the model’s cost per 1,000 requests by its quality score and you have a cost per quality point you can compare directly across candidates and across test sets of different sizes. Run this as a short script that swaps one model string and reads the per-request cost from the usage object in each response. We return the usage object in every non-streaming response and in the final message of every streaming response without any extra request parameter, as described in the usage accounting docs . usage.cost is the amount we charged your account for that request, in USD. The OpenAI Python SDK keeps response fields it doesn’t define, so r.usage.cost is available as an attribute.
import json
import os
from openai import OpenAI
client = OpenAI(
base_url = "https://openrouter.ai/api/v1" ,
api_key = os.environ[ "OPENROUTER_API_KEY" ],
)
ROUTING_PROMPT = open ( "routing_prompt.txt" ).read()
test_set = json.load( open ( "test_set.json" )) # 20-50 of your own {"ticket": ..., "label": ...} examples
candidates = [
"openai/gpt-5.6-luna" ,
"google/gemini-3.7-flash" ,
"anthropic/claude-opus-5" ,
]
for model in candidates:
total_cost, correct = 0.0 , 0
for ex in test_set:
r = client.chat.completions.create(
model = model,
messages = [
{ "role" : "system" , "content" : ROUTING_PROMPT },
{ "role" : "user" , "content" : ex[ "ticket" ]},
],
)
total_cost += r.usage.cost
answer = (r.choices[ 0 ].message.content or "" ).strip()
correct += int (answer == ex[ "label" ])
score = correct / len (test_set)
points = score * 100
cost_per_1k = total_cost / len (test_set) * 1000
cost_per_point = cost_per_1k / points if points else float ( "inf" )
print ( f " { model } : score= { score :.0% } cost/1K=$ { cost_per_1k :.2f } cost/point=$ { cost_per_point :.4f } " )
A response can carry no text content, for example when the model refuses, so the script treats missing content as an empty answer, counts it as incorrect, and keeps its cost in the total. A model that scores zero has no cost per point, so the script prints inf for it instead of dividing by zero.
Here is a worked example for a support-routing task. Assume each request sends about 700 input tokens, the ticket plus a short system prompt, and returns about 150 output tokens. Cost per 1,000 requests is (700 / 1M × input price) + (150 / 1M × output price) , multiplied by 1,000. Cost per quality point is that cost divided by the quality score.
The prices are the listed rates in our model catalog on 2026-09-18 for GPT-5.6 Luna , Gemini 3.7 Flash , and Claude Opus 5 . They are the model-level prices the catalog lists. Individual provider endpoints and the flex and priority service tiers bill at their own rates, and GPT-5.6 Luna bills a higher rate on requests with 272,000 or more prompt tokens. The token counts and quality scores are illustrative. Your own run replaces them.
Model Price in / out (per M tokens) Cost / 1K requests Quality score Cost / quality point Clears 85% bar?
openai/gpt-5.6-luna $0.20 / $1.20 $0.32 82 $0.0039 No (−3)
google/gemini-3.7-flash $0.75 / $3.75 $1.09 89 $0.0122 Yes (+4)
anthropic/claude-opus-5 $5.00 / $25.00 $7.25 97 $0.0747 Yes (+12)
Read the table in the order the framework runs. First gate on the bar. At 85%, GPT-5.6 Luna is out at 82, so its low cost per quality point doesn’t matter. A model below the bar is disqualified no matter how cheap it is per point. That leaves Gemini 3.7 Flash and Claude Opus 5.
Between the two that clear, take the cheaper one. Gemini 3.7 Flash costs $1.09 per 1,000 requests. Claude Opus 5 costs $7.25 for a score the task didn’t require. Gemini 3.7 Flash is the pick, and it clears with a 4-point margin.
Change the situation and the answer changes. If a misroute means a missed SLA and you set the bar at 95%, only Claude Opus 5 clears and you pay the $7.25. The goal is not to pick the highest score. It is to know which model clears your bar for the least money, and to see the gap you pay for when you reach past it.
Step 3: Pick the cheapest model that clears the bar with margin
Pick the cheapest model that clears the bar with margin, not the model that wins outright. The margin matters because the numbers move.
Model scores drift as providers update weights, your own traffic shifts over time, and new versions change the picture. Set the margin from observation rather than picking a fixed percentage. Run the candidates more than once, or across a fresh slice of traffic, record how much the score moves between runs, and require the winner to clear the bar by more than that observed swing.
A model that only reaches the bar on a single small sample should have to clear a higher internal target before you commit to it, so that ordinary variation does not push it below the bar in production.
Applied across a few tasks, the framework looks like this. The cost figures use the same listed prices as the table above at the input and output token sizes noted per row. The quality bars are examples, and the cost you record comes from usage.cost .
Task Quality bar Cheapest tier that clears it Cost / 1K requests
Support triage (700 in / 150 out tokens) 80% Cheap, openai/gpt-5.6-luna $0.32
Code review (4,000 in / 1,000 out tokens) 88% Mid-tier, google/gemini-3.7-flash $6.75
Compliance review (3,000 in / 600 out tokens) 95% Frontier, anthropic/claude-opus-5 $30.00
The method is the same in each row, and the answer differs because the bar differs. The measured tradeoff was true on the day you measured it. Recheck it when a newer model is released, because the model that cleared your bar yesterday may now cost the same as one that sits just below it.
Recheck the comparison when models or prices change
New models are released often, and prices on existing models change. Between May and September 2026, four versions of Gemini Flash were added to our catalog: Gemini 3.5 Flash on 2026-05-19, Gemini 3.6 Flash on 2026-07-21, Gemini 3.7 Flash on 2026-08-13, and Gemini 3.8 Flash on 2026-09-02. A comparison you ran a few months ago can already be out of date, so rerun it whenever a candidate is updated or its pricing changes.
Rerunning is cheap on our side. We expose one OpenAI-compatible API, so swapping a model is a config change. The base URL, key, and SDK stay the same. You change the model string from openai/gpt-5.6-luna to anthropic/claude-opus-5 and the script above runs against the new model.
You don’t write a new integration for each provider you want to test, which is what makes it practical to rerun your examples against a new release the week it comes out. That same one-line swap is what model routing in a multi-agent setup relies on.
Two things keep the framework grounded in real numbers. We show per-model pricing in the model catalog before you commit, so the cost side of your comparison starts from live data. And because every response reports usage.cost , the spend you measured is what we billed, not an estimate from a rate card.
Frequently asked questions
What is the cost versus quality tradeoff for AI models?
It is the tradeoff between what a request costs and how well a model performs on your task. Speed is a third constraint. A model that is cheap and accurate but too slow for the task is still out. Choosing a model means deciding how much quality the task needs and paying for that, rather than paying for the highest score available.
What is the best model for AI agents?
There is no single best model. The right model for an agent depends on the quality bar its task has to clear, the latency it has to respond within, and the volume it has to handle. A model that clears the bar for compliance review costs more than a task like support triage needs, and a model that is cheap enough for triage may not clear the bar for compliance review.
How are AI models priced for agent workloads?
Text models in our catalog list a prompt rate and a completion rate per token. Some endpoints also list rates for cached input, reasoning tokens, per-request charges, and flex or priority service tiers. Agent workloads add tool calls and retries on top of the first request, so the token price of one completion understates the cost of one task. Measure cost across a full agent run and read it from usage.cost in each response.
Is a multi-agent design more or less expensive than a single agent?
It depends on how the work splits. A multi-agent design can send routing and classification to a cheap model and call a more expensive model only for the requests that need it, which costs less than sending every request to the expensive model. It also adds requests, so measure the full run rather than assuming the split saves money.
Conclusion
Define the bar, measure complete runs on your own examples, and choose the lowest-cost model that clears the bar by more than the swing you observed between runs. Keep the script and the example set so you can rerun the comparison when models, prices, or your traffic change.
References
首次收录 · 2026-10-02 · 9.95 分