今天,我们推出了新的模型路由器基准测试页面,用于在日益增多的路由器选项之间比较性能。每款路由器都经过一套来自不同领域的基准测试,以评估其在质量、速度和成本方面的表现。
代表最佳成本效益和智能水平的单个模型与路由器一同展示,作为比较的基线。请持续关注此页面,因为我们将不断添加新的路由器,并随着路由器的改进更新评分。
OpenRouter 推出了首批两款模型路由器:我们的 Auto Router 和 Free Models Router。它们的创建旨在为您提供一个可靠的选项,始终跟上最新模型的步伐,同时符合您偏好的成本水平。此后,我们持续改进了它们,包括利用市场智慧来指导路由决策。
在最近几个月里,许多路由器相继推出,它们利用模型生态系统中不同的优势来优化性能。
成本差异:例如,将 DeepSeek v4 Flash 与 GPT-6 Astra 进行比较,GPT-6 Astra 每令牌的平均支付价格高出 48 倍以上,在 Codex 中运行一次平均 10 到 49 轮会话的成本更是其 21 倍。路由器可以将任务的部分切换至成本较低模型,从而节省费用。
任务多样性:例如,在法律研究方面表现最佳的模型,未必是编码方面的最佳选择。路由器可以在它们之间切换,以利用每个模型的优势。
会话阶段:智能体会话的不同部分可能需要不同水平的智能。路由器可以在模型或推理努力之间切换,以适应会话中需要复杂规划或常规执行的部分。
在实践中,这些技术各自面临挑战,有时会导致性能不如直接使用单个模型:
在模型之间切换意味着请求成本会更高,因为输入缓存需要重建。
路由器通常没有足够的信息来仅通过查看提示词来评估任务复杂度。
用于判断会话阶段的启发式信号可能与模型的下一步行动不一致。
您在路由决策上进行的处理越多,引入的延迟就越大。
我们提供这些模型路由器基准测试,以展示哪些路由器克服了这些固有的挑战,从而帮助您决定何时(以及是否)适合使用它们。它们旨在帮助您找到符合您所需性能特征的路由器和模型组合。
路由器类型
“路由器”一词被互换使用,指代许多不同的事物。让我们将其分解。
提供商路由:模型通常由多个推理提供商提供服务,因此提供商路由器会选择符合您对价格、速度、正常运行时间和数据政策偏好的提供商,如果失败则回退到另一个。这是 OpenRouter 自第一天起就一直在做的事情。
模型路由:您不必直接选择单个模型,而是可以将请求发送到 Auto Router 等路由器,由它决定哪个模型应进行回答。它们的行为就像模型一样,但可能使用一个或多个模型来响应您的请求。
我们基准测试的模型路由器采取了以下几种不同的方法:
执行混合模型的 routers:Unbiased 的 Pareto 和 Sakana 的 Fugu 是推理提供商,它们使用混合模型,有时采用前沿升级策略。它们不披露所使用的模型或切换频率,并按标准化的每令牌费率计费。
选择单个模型的路由器:Auto Router 和 Jev Router 为每一轮选择一个模型。模型选择是透明的,您将根据所选模型的标准费率计费。
在预定义模型集之间切换的路由器:NVIDIA 的 Switchyard 将一个较便宜的模型与一个更强的模型配对,并在智能体处理任务时在它们之间切换。基准测试了多个配对以展示哪些模型组合效果最佳,但您可以使用任何您想要的配对来运行它。
还有其他几种风格,例如 Pareto Code 或 -latest 模型标识符,它们都指向单一模型。Pareto Code 会根据你所需的智能水平选择成本最优的模型,而 -latest 标识符则指向某个模型家族中最新的模型。Fusion 是另一种风格,它会将相同的请求发送给一个模型委员会,并综合出最佳答案。这些风格未经过基准测试,因为它们服务于特定的目的,需要不同的比较方法。
Router Index 融合质量、速度与成本
基准测试揭示了多个维度的性能表现,使得难以判断哪个路由器的表现最佳。Router Index 将质量、速度和成本综合为每个基准测试的 0 到 10 分。默认情况下,它将质量(基准测试得分)的权重设为 60%,任务耗时占 20%,成本占 20%。你的优先级可能有所不同,因此你可以通过页面上的滑块调整索引权重。
只需替换模型,即可尝试上述任意路由器
与任何基准测试一样,我们的 Model Router Benchmarks 仅代表通用任务,而非你自己的具体工作。列出的所有路由器均可在 OpenRouter 上使用。你可以在任何应用、框架或通过我们的 API 中将其替换为你的模型。
Today we’re introducing a new Model Router Benchmarks page for comparing performance across the growing number of router options. Every router was run through a set of benchmarks from varied domains to score them on quality, speed, and cost.
Individual models representing both best-in-class cost efficiency and intelligence are shown alongside routers as a baseline for comparison. Keep an eye on this page over time, as we’ll continually add new routers and update scores as the routers improve.
OpenRouter introduced two of the first model routers with our Auto Router and Free Models Router . They were created so you’d have a dependable option that stays up to date with the latest models, at your preferred level of cost . Since then we’ve continually improved them, including by using the wisdom of the market to inform routing decisions.
In recent months, many routers launched that use the differing strengths across the model landscape to optimize performance.
Cost differences: For example, comparing DeepSeek v4 Flash with GPT-6 Astra, GPT-6 Astra’s average price paid per token is over 48x higher, and it’s 21x more expensive to run an average 10 to 49 turn session in Codex. Routers can switch parts of a task to a lower-cost model to save on cost.
Task diversity: For example, the model that performs best on legal research may not be the same as the model that’s best for coding. Routers can switch between them to leverage each model’s strengths.
Session stage: Different parts of an agentic session may require different levels of intelligence. Routers can switch between models or reasoning efforts to adjust for parts of a session that require complex planning or mundane execution.
In practice, these techniques each have challenges that can sometimes lead to worse performance than using a single model directly:
Switching between models means requests get more expensive while the input cache is rebuilt.
Routers often don’t have enough information to assess task complexity by simply looking at the prompt.
Heuristic signals for session stage may not agree with the model’s next step.
The more processing you run on the routing decisions, the more latency is introduced.
We offer these Model Router Benchmarks to show which routers overcome these inherent challenges so you can decide when (and if) it’s the right time to put them to use. They’re designed to help you find the router and model combination that matches the performance characteristics you need.
Types of routers
The term “router” is used interchangeably to mean many different things. Let’s break it down.
Provider routing: Models are usually served by several inference providers, so a provider router picks the one that fits your preferences for price, speed, uptime, and data policy, then falls back to another if it fails. This is what OpenRouter has done since day one.
Model routing: Instead of choosing a single model directly, you can send a request to a router such as the Auto Router , and it decides which model should answer. These behave just like models, but may use one or more models to respond to your request.
The model routers we benchmarked take a few different approaches:
Routers that execute a blend of models: Unbiased’s Pareto and Sakana’s Fugu are inference providers that use a blend of models, sometimes with frontier escalation. They don’t disclose the models used or how often switching happens, and they bill at standardized per-token rates.
Routers that select a model: Auto Router and Jev Router select a model for each turn. The model selection is transparent, and you’re billed at the standard rate for whichever model was selected.
Routers that swap between a pre-defined set of models: NVIDIA’s Switchyard pairs a cheaper model with a stronger one and switches between them as an agent works through a task. Several pairs were benchmarked to show which models work best together, but you can run it with any pairing you want.
There are a couple of other styles, such as Pareto Code or the -latest model slugs , that alias to a single model. Pareto Code picks the most cost-optimal model for your desired level of intelligence, and the -latest slugs point to the most recent model within a family. Fusion is another style that runs the same request across a council of models and synthesizes the best answer. These styles weren’t benchmarked because they serve specialized purposes that warrant different comparison methods.
Router Index blends quality, speed, and cost
Benchmarks reveal performance across several dimensions, making it hard to tell which router performed best. The Router Index synthesizes quality, speed, and cost into a 0 to 10 score for each benchmark. By default, it weights quality (benchmark score) at 60%, time per task at 20%, and cost at 20%. Your priorities may differ, so you can drag a slider on the page to adjust the index weights.
Try any of these routers by simply swapping them in for your model
As with any benchmark, our Model Router Benchmarks are only representative of general tasks rather than your own work. All of the routers listed are available on OpenRouter. Swap them in for your model in any app, harness, or through our API.
| 刊期 | 得分 | 排名 | 结果 |
|---|---|---|---|
| 2026-10-05 | 8.77 | 16 | 入选 |
| 2026-10-04 | 9.22 | 23 | 未入选 |
| 2026-10-03 | 9.95 | 29 | 未入选 |