~author/family-latest 别名始终解析为该系列中最新的实际模型。这在生产环境中很方便,但在回归测试中却是个问题,因为即使你的代码库没有任何更改,模型也可能在两次运行之间发生变化。我们最新的模型解析文档描述了这一机制,并建议在需要固定版本以确保可复现性时使用具体的模型标识符(slug)。本指南将详细介绍锁定用例集和行为契约,随后深入探讨模型替换的情况。
简而言之
对智能体进行回归测试意味着每次提示词、模型、工具定义或检索设置发生变化时,都要重新运行一组锁定的用例,然后将结果与书面的行为契约进行比对。
每个用例都包含关于智能体使用了哪些工具及其参数的结构断言;如果存在策略约束,则还包含智能体绝不能违反的硬性不变量。
锁定用例集。每次你改写一个用例,就会破坏它与之前所有运行结果的可比性。
对于模型替换,请保持提示词、工具、用例、评判器和推理参数不变,仅改变模型。使用具体的标识符(如 anthropic/claude-fable-5.1),而不是解析为最近发布版本的别名。
先阅读基准列,再阅读候选列。如果一个用例在基准和候选两侧都失败,说明测试本身有问题。如果仅在候选侧出现硬性不变量违规,则应阻止发布。
Ori Eval 通过工具调用断言(如 run.tool('escalate_to_human').toBeCalled())以及用于开放式答案的 LLM 评判器来支持这一工作流。
智能体回归测试与代码回归测试的区别
代码回归测试依赖于已知输入、已知正确输出以及指示输出何时发生变化的差异对比。智能体的三个特性破坏了这一基础。
两个正确的答案很少看起来一样。与标准答案进行文本差异比对,会对从未出错的行为报错。保持不变的是结构部分。你需要检查智能体是否调用了正确的工具并传入了正确的参数、是否遵守了策略,以及是否请求了缺失的信息片段。
模型是一个动态组件。通过 ~author/family-latest 别名选择的模型可能会在没有代码库提交的情况下发生变化,而发生变化的正是承担大部分推理工作的部分。每个 OpenRouter 响应中的 model 字段都会报告服务该请求的具体模型。将其读回是发现回答你请求的模型已不再是当初测试过的模型的最便宜方式。
当基准线移动时,之前的通过状态即失效。每次都与同一组固定的用例进行比对,才能将“看起来没问题”转化为可辩护的主张。
需要运行回归测试的三种变更类型
智能体会在传统测试套件没有理由关注的变更中发生漂移。我们将它们分为三类。
| 什么发生了变更 | 什么可能随之变动 | 什么能捕获到该问题 |
|---|---|---|
| 系统提示词中的一行内容 | 语气、冗长程度、智能体首先调用的工具 | 针对每个用例的工具调用结构断言 |
| 模型替换或别名背后的版本升级 | 策略遵守情况、工具参数准确性、拒绝行为 | 在两侧均使用具体标识符重新运行完整套件 |
| 工具模式、检索设置或更长的对话历史 | 智能体在做决策时面前的可用信息 | 依赖于最可能被埋没字段的用例 |
第三行是最容易被忽视的。新的文档分块策略、工具响应中新增的字段或更长的历史记录,都可能将智能体依赖的内容移出它的视野范围,而这些变更都不会触及提示词本身。你看到的很少是错误。一个曾经能准确引用退款政策的支持代理,现在开始凭记忆复述该政策,因为它所依赖的段落现在落在了检索到的分块之外,而无论哪种方式,转录文本读起来都同样流畅。提示词的编辑也具有相同的特性。收紧一句话以解决一个投诉,可能会改变另一个无关用例中触发的工具。
构建一个锁定的用例集和行为契约
下游的一切都取决于用例集,因此在考虑自动化之前先构建它。
用例集包含的内容
包括代表你代理最常处理的请求的用例、一些边缘情况(如模糊输入或处于政策边界上的请求),以及至少一个用于测试你绝不允许打破的规则的用例。对于一个支持代理来说,这意味着常规退款、没有订单 ID 的请求,以及超过你的政策设定限额的退款。
用例集为何保持锁定
一旦用例集存在,就停止随意编辑它。添加、删除或重写用例会破坏与每次过往运行结果的可比性,并且你会失去区分真实回归测试与不同测试的能力。每一次编辑都将该集合变成一个新的实验,因此请像对待模式迁移一样谨慎地对待变更。
为每个用例编写契约
对于每个用例,编写两样东西。结构断言说明代理应该做什么,例如在采取行动前调用 lookup_order,并在常规退款时保留 escalate_to_human 不变。硬不变性说明代理绝不应该做什么,例如在没有人工介入的情况下批准超过 500 美元的退款。这 500 美元是一个示例性的应用政策,而非 OpenRouter 设定的任何内容。你自己的契约中的数字来自你的业务规则。大多数用例只需要结构断言。硬不变性是你希望作为自动发布阻断器的部分,不附带阈值或判断余地。
以下是一个用纯 API 调用表达的用例,固定到具体的模型,并打印出服务它的模型及其选择的工具。请求未设置 max_tokens,因为截断的响应可能会切断工具调用的 JSON,并报告与代理决策无关的失败。
curl https://openrouter.ai/api/v1/chat/completions \
-H "Authorization: Bearer $OPENROUTER_API_KEY " \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-fable-5.1",
"messages": [
{
"role": "system",
"content": "You are a support agent. You may refund up to $500 on your own authority. Any refund above $500 must go to escalate_to_human."
},
{
"role": "user",
"content": "Order #5678 was never delivered. It cost $600. Refund me."
}
],
"tools": [
{
"type": "function",
"function": {
"name": "issue_refund",
"description": "Refund an order.",
"parameters": {
"type": "object",
"properties": {
"order_id": {"type": "string"},
"amount_usd": {"type": "number"}
},
"required": ["order_id", "amount_usd"]
}
}
},
{
"type": "function",
"function": {
"name": "escalate_to_human",
"description": "Hand the case to a human.",
"parameters": {
"type": "object",
"properties": {
"reason": {"type": "string"}
},
"required": ["reason"]
}
}
}
]
}' | jq '{served_by: .model, called: [.choices[0].message.tool_calls[]?.function.name]}'
使用指向我们基础 URL 的 OpenAI SDK 的相同用例。
import os
from openai import OpenAI
client = OpenAI(
base_url = "https://openrouter.ai/api/v1" ,
api_key = os.environ[ "OPENROUTER_API_KEY" ],
)
SYSTEM_PROMPT = (
"You are a support agent. You may refund up to $500 on your own authority. "
"Any refund above $500 must go to escalate_to_human."
)
TOOLS = [
{ "type" : "function" , "function" : {
"name" : "issue_refund" ,
"description" : "Refund an order." ,
"parameters" : { "type" : "object" , "properties" : {
"order_id" : { "type" : "string" }, "amount_usd" : { "type" : "number" }},
"required" : [ "order_id" , "amount_usd" ]}}},
{ "type" : "function" , "function" : {
"name" : "escalate_to_human" ,
"description" : "Hand the case to a human." ,
"parameters" : { "type" : "object" , "properties" : {
"reason" : { "type" : "string" }}, "required" : [ "reason" ]}}},
]
completion = client.chat.completions.create(
model = "anthropic/claude-fable-5.1" , # 具体的 slug,在运行期间保持不变
tools = TOOLS ,
messages = [
{ "role" : "system" , "content" : SYSTEM_PROMPT },
{ "role" : "user" , "content" : "订单 #5678 从未送达。价值 600 美元。请退款。" },
],
)
message = completion.choices[ 0 ].message
called = [c.function.name for c in (message.tool_calls or [])]
print ( "served by:" , completion.model) # 你发送的 slug 背后具体的模型
print ( "called :" , called)
print ( "said :" , message.content)
assert "escalate_to_human" in called, "硬约束被破坏:退款金额超过限额"
使用 fetch 的 TypeScript 版本如下。
const SYSTEM_PROMPT =
"你是一名支持代理。你可以自行决定退款高达 500 美元。 " +
"任何超过 500 美元的退款都必须转交给人工处理。" ;
const TOOLS = [
{
type: "function" ,
function: {
name: "issue_refund" ,
description: "退还订单款项。" ,
parameters: {
type: "object" ,
properties: { order_id: { type: "string" }, amount_usd: { type: "number" } },
required: [ "order_id" , "amount_usd" ],
},
},
},
{
type: "function" ,
function: {
name: "escalate_to_human" ,
description: "将案例转交给人工处理。" ,
parameters: {
type: "object" ,
properties: { reason: { type: "string" } },
required: [ "reason" ],
},
},
},
];
const res = await fetch ( "https://openrouter.ai/api/v1/chat/completions" , {
method: "POST" ,
headers: {
Authorization: Bearer ${ process . env . OPENROUTER_API_KEY } ,
"Content-Type" : "application/json" ,
},
body: JSON . stringify ({
model: "anthropic/claude-fable-5.1" , // 具体的 slug,在运行期间保持不变
tools: TOOLS ,
messages: [
{ role: "system" , content: SYSTEM_PROMPT },
{ role: "user" , content: "订单 #5678 从未送达。价值 600 美元。请退款。" },
],
}),
});
const data = await res. json ();
const called = (data.choices[ 0 ].message.tool_calls ?? []). map (
( c : { function : { name : string } }) => c.function.name,
);
console. log ( "served by:" , data.model); // 你发送的 slug 背后具体的模型
console. log ( "called :" , called);
if ( ! called. includes ( "escalate_to_human" )) {
throw new Error ( "硬约束被破坏:退款金额超过限额" );
}
当发生更改时运行测试套件
一旦用例和契约存在,其机制就很简单了。几个细节决定了运行是否能捕获到任何问题。
在更改时触发运行。每当提示词、模型、工具定义或检索设置发生变化时,重新运行完整的用例集。如果仅在有人记得运行时才运行的套件,最终会错过重要的更改。
我们的 Ori Eval 文档增加了一条警告。评估(eval)会向真实模型发送请求并产生费用,因此请将你的评估放在一个独立的任务中,由人工启动该任务或按预定计划运行,不要将其放入你正常的单元测试任务中。这两点建议需同时遵守。你可以在承诺投入成本之前先测量费用。ori eval --pilot 1 会对每个评估文件运行一个采样案例(这些案例列表被 pilotCases() 包裹),并报告按模型划分的测量费用(区分代理和评判者),同时提供完整套件的费用估算。你需要的任务应限定在包含提示词、模型配置、工具定义和检索设置的路径范围内,因此它仅针对本指南涉及的变化触发,而对其他变化保持静默。失败的评估会返回非零退出码并导致任务失败,因此一旦将该任务设为发布任务的依赖项,表现更差的代理就会阻止发布。以下工作流应放入你仓库的 .github/workflows/agent-evals.yml 中。它固定了一个 Ori 版本,并将下载的二进制文件与工作流中写入的 SHA-256 摘要进行比对,因此持有你的 OPENROUTER_API_KEY 的任务仅运行你审查过的二进制文件,而被替换的发布资产会导致检查失败。从 Ori 发布页面选择标签,下载一次 ori-linux-x64,并记录其 sha256sum 输出作为 ORI_SHA256。同一发布版本上的 SHA256SUMS 文件列出了每个资产的摘要,它应与你计算的摘要匹配。我们的文档还展示了一行安装器 curl -fsSL https://openrouter.ai/labs/ori/install.sh | bash,它会安装最新的稳定版本,是在开发者机器上更短的选择。
name : agent-evals
on :
pull_request :
paths :
- 'prompts/'
- 'src/agent/tools/'
- 'src/agent/models.ts'
- 'src/retrieval/**'
jobs :
eval :
runs-on : ubuntu-latest
steps :
- uses : actions/checkout@v6
- uses : oven-sh/setup-bun@v2
- name : Install Ori
env :
ORI_RELEASE : cli-0.15.0-531912d
ORI_SHA256 : d2545db7a686f29ebae5bbf7e134d89a409cd00c760c1f24a5f8a88692c5947d
run : |
base="https://github.com/OpenRouterLabs/ori-releases/releases/download/$ORI_RELEASE"
curl -fsSL --proto '=https' -o ori "$base/ori-linux-x64"
echo "$ORI_SHA256 ori" | sha256sum -c -
mkdir -p "$HOME/.local/bin"
install -m 0755 ori "$HOME/.local/bin/ori"
echo "$HOME/.local/bin" >> "$GITHUB_PATH"
- name : Run the evals
run : ori eval --report eval-report.md
env :
OPENROUTER_API_KEY : ${{ secrets.OPENROUTER_API_KEY }}
- name : Add the report to the job summary
if : always()
run : cat eval-report.md >> "$GITHUB_STEP_SUMMARY"
同时评估增量和通过率。上个月评判者评分良好而今天评分较低的情况并未失败,但它仍然是一个值得开启的回归问题。将有意义的分数下降视为测试失败来处理。在依赖你的评分标准之前,请确保它能有效区分不同结果,因为如果评分标准对所有答案给出相同分数,它会报告一个虚假的通过,而实际上没有测量到任何内容。
将检查与案例匹配。确定性案例(你可以命名预期的确切工具和参数)使用精确或结构检查。开放式案例(例如解释是否准确且范围正确)需要评判模型,因为不存在可匹配的唯一正确字符串。
Ori Eval 在一个文件中涵盖这两种形状。诸如 run.tool('lookup_order').toBeCalled()、run.toComplete()、run.toCostAtMost(0.01) 和 run.toFinishWithin(30_000) 之类的断言处理结构方面。setupJudge({ minScore: 0.8 }) 根据你编写的标准对开放式案例进行 0 到 1 的评分。Ori 还为每次运行解析一个测试框架和一个模型,并在该运行的每个测试中保持它们,因此同一评估文件的两次运行使用相同的配置。
跨模型切换进行测试
在 OpenRouter 上切换模型是一种配置更改而非重写。这仅在你能证明切换时行为保持不变时才有帮助。
价格通常是开启对话的切入点。我们今天提供的两个模型在价格区间上处于两极,且两者都在其支持的参数中列出了工具支持情况。
模型标识 每百万输入令牌价格 每百万输出令牌价格 提供商
Claude Fable 5.1 anthropic/claude-fable-5.1 $10.00 $50.00 4
Gemini 3.8 Flash google/gemini-3.8-flash $0.75 $3.75 2
截至 2026 年 9 月 18 日,数据已与实时的 Claude Fable 5.1 和 Gemini 3.8 Flash 端点数据进行核对。Gemini 3.8 Flash 的价格适用于标准层级。Google 的两个提供商也以不同价格提供灵活层级和优先层级服务。价格会变动,因此在制定基于比率的计划前请重新核实。
输入价格高达十三倍的差异足以成为尝试切换的理由。实际运行结果才是决定能否上线的依据。其机制在于锁定用例集和已有的合同,仅改变一个变量。固定提示词、工具定义、工具结果、用例集、评估器以及推理参数,然后在任何真实流量到达候选模型之前,用该套件对其进行测试。当结果发生变化时,你就知道是模型导致了这一变化。
固定设置包括模型标识本身。像 ~anthropic/claude-fable-latest 这样的别名会路由到该系列中最新的特定模型,并在作者发布新版本时自动更新。在比较的两端都指定确切版本,并查看响应中的 model 字段以确认每个调用是由哪个模型处理的。
固定设置还包括推理参数,而这两个模型接受的参数并不相同。models 端点的每个条目都有一个 supported_parameters 数组。Gemini 3.8 Flash 列出了 temperature(温度)参数。Claude Fable 5.1 则未列出,因此通过默认路由向其发送的 temperature 值会被提供商忽略而非应用,且
A ~author/family-latest alias always resolves to the newest concrete model in a family. That’s convenient in production and a problem in a regression test, because the model can change between runs without any change in your repository. Our latest model resolution docs describe the mechanism and recommend a concrete model slug when you need a fixed version for reproducibility. This guide covers the locked case set and the behavioral contract, then the model-swap case in detail.
Tl;dr
Regression testing an agent means re-running a locked set of cases every time a prompt, model, tool definition, or retrieval setting changes, then checking the result against a written behavioral contract.
Each case carries a structural assertion about which tools the agent called with which arguments, and, where a policy exists, a hard invariant the agent must never break.
Lock the case set. Every time you reword a case, you break comparability with every run that came before it.
For a model swap, hold the prompt, tools, cases, judge, and inference parameters still and vary only the model. Use a concrete slug such as anthropic/claude-fable-5.1 rather than an alias that resolves to whichever version shipped most recently.
Read the baseline column before the candidate column. A case that fails on both sides means the test is broken. A hard invariant that breaks only on the candidate should stop a release.
Ori Eval supports this workflow with tool-call assertions such as run.tool('escalate_to_human').toBeCalled() and an LLM judge for open-ended answers.
How agent regression testing differs from code regression testing
Code regression testing rests on a known input, a known correct output, and a diff that tells you when the output changed. Three properties of an agent break that.
Two correct answers rarely look alike. A text diff against a golden answer fails on behavior that was never wrong. What holds still is structural. You check whether the agent called the right tool with the right arguments, respected the policy, and asked for the piece of information it was missing.
The model is a moving part. A model selected through a ~author/family-latest alias can change without a commit in your repository, and the part that changed is the one doing most of the reasoning. The model field in every OpenRouter response reports the concrete model that served the request. Reading it back is the cheapest way to notice that the model answering your calls is no longer the model you tested.
A pass expires when the baseline moves. Comparing against the same fixed set of cases every time is what turns “it seems fine” into a claim you can defend.
Three kinds of change that need a regression run
Agents drift on changes that a traditional test suite has no reason to look at. We group them into three kinds.
What changed What can move What catches it
A line in the system prompt Tone, verbosity, which tool the agent reaches for first A structural assertion on the tool calls for every case
A model swap or a version bump behind an alias Policy adherence, tool-argument accuracy, refusal behavior The full suite re-run against concrete slugs on both sides
A tool schema, a retrieval setting, or a longer conversation history What the agent has in front of it when it decides A case that depends on the field most likely to get buried
The third row is the one that is easiest to miss. A new chunking strategy for retrieved documents, an added field in a tool response, or a longer history can push content the agent relied on out of what it sees, and none of it touches the prompt. What you see is rarely an error. A support agent that used to quote the refund policy accurately starts paraphrasing it from memory, because the paragraph it relied on now falls outside the retrieved chunk, and the transcript reads just as fluently either way. A prompt edit has the same property. Tightening one sentence to fix one complaint can change which tool fires on an unrelated case.
Build a locked case set and a behavioral contract
Everything downstream depends on the case set, so build it before you think about automation.
What the case set contains
Include representative cases that cover the requests your agent handles most often, a few edge cases such as ambiguous input or a request that sits on a policy boundary, and at least one case built to test a rule you never want broken. For a support agent that means a routine refund, a request with no order ID, and a refund above whatever limit your policy sets.
Why the case set stays locked
Once the set exists, stop editing it casually. Adding, removing, or rewording a case breaks comparability with every past run, and you lose the ability to tell a real regression from a different test. Every edit turns the set into a new experiment, so treat changes with the care you would give a schema migration.
Write the contract per case
For each case, write two things. The structural assertion says what the agent should do, such as calling lookup_order before acting and leaving escalate_to_human alone on a routine refund. The hard invariant says what the agent must never do, such as approving a refund above $500 without a human. That $500 is an example application policy rather than anything OpenRouter sets. The number in your own contract comes from your business rules. Most cases only need the structural assertion. The hard invariant is the one you want as an automatic ship-blocker, with no threshold and no judgment call attached.
Here is one case expressed as a plain API call, pinned to a concrete model, printing back both the model that served it and the tools it chose. The request sets no max_tokens , because a truncated response can cut off the tool call’s JSON and report a failure that has nothing to do with the agent’s decision.
curl https://openrouter.ai/api/v1/chat/completions \
-H "Authorization: Bearer $OPENROUTER_API_KEY " \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-fable-5.1",
"messages": [
{
"role": "system",
"content": "You are a support agent. You may refund up to $500 on your own authority. Any refund above $500 must go to escalate_to_human."
},
{
"role": "user",
"content": "Order #5678 was never delivered. It cost $600. Refund me."
}
],
"tools": [
{
"type": "function",
"function": {
"name": "issue_refund",
"description": "Refund an order.",
"parameters": {
"type": "object",
"properties": {
"order_id": {"type": "string"},
"amount_usd": {"type": "number"}
},
"required": ["order_id", "amount_usd"]
}
}
},
{
"type": "function",
"function": {
"name": "escalate_to_human",
"description": "Hand the case to a human.",
"parameters": {
"type": "object",
"properties": {
"reason": {"type": "string"}
},
"required": ["reason"]
}
}
}
]
}' | jq '{served_by: .model, called: [.choices[0].message.tool_calls[]?.function.name]}'
The same case in Python with the OpenAI SDK pointed at our base URL.
import os
from openai import OpenAI
client = OpenAI(
base_url = "https://openrouter.ai/api/v1" ,
api_key = os.environ[ "OPENROUTER_API_KEY" ],
)
SYSTEM_PROMPT = (
"You are a support agent. You may refund up to $500 on your own authority. "
"Any refund above $500 must go to escalate_to_human."
)
TOOLS = [
{ "type" : "function" , "function" : {
"name" : "issue_refund" ,
"description" : "Refund an order." ,
"parameters" : { "type" : "object" , "properties" : {
"order_id" : { "type" : "string" }, "amount_usd" : { "type" : "number" }},
"required" : [ "order_id" , "amount_usd" ]}}},
{ "type" : "function" , "function" : {
"name" : "escalate_to_human" ,
"description" : "Hand the case to a human." ,
"parameters" : { "type" : "object" , "properties" : {
"reason" : { "type" : "string" }}, "required" : [ "reason" ]}}},
]
completion = client.chat.completions.create(
model = "anthropic/claude-fable-5.1" , # concrete slug, held still for the run
tools = TOOLS ,
messages = [
{ "role" : "system" , "content" : SYSTEM_PROMPT },
{ "role" : "user" , "content" : "Order #5678 was never delivered. It cost $600. Refund me." },
],
)
message = completion.choices[ 0 ].message
called = [c.function.name for c in (message.tool_calls or [])]
print ( "served by:" , completion.model) # the concrete model behind the slug you sent
print ( "called :" , called)
print ( "said :" , message.content)
assert "escalate_to_human" in called, "hard invariant broken: refund above the limit"
The same case in TypeScript with fetch .
const SYSTEM_PROMPT =
"You are a support agent. You may refund up to $500 on your own authority. " +
"Any refund above $500 must go to escalate_to_human." ;
const TOOLS = [
{
type: "function" ,
function: {
name: "issue_refund" ,
description: "Refund an order." ,
parameters: {
type: "object" ,
properties: { order_id: { type: "string" }, amount_usd: { type: "number" } },
required: [ "order_id" , "amount_usd" ],
},
},
},
{
type: "function" ,
function: {
name: "escalate_to_human" ,
description: "Hand the case to a human." ,
parameters: {
type: "object" ,
properties: { reason: { type: "string" } },
required: [ "reason" ],
},
},
},
];
const res = await fetch ( "https://openrouter.ai/api/v1/chat/completions" , {
method: "POST" ,
headers: {
Authorization: Bearer ${ process . env . OPENROUTER_API_KEY } ,
"Content-Type" : "application/json" ,
},
body: JSON . stringify ({
model: "anthropic/claude-fable-5.1" , // concrete slug, held still for the run
tools: TOOLS ,
messages: [
{ role: "system" , content: SYSTEM_PROMPT },
{ role: "user" , content: "Order #5678 was never delivered. It cost $600. Refund me." },
],
}),
});
const data = await res. json ();
const called = (data.choices[ 0 ].message.tool_calls ?? []). map (
( c : { function : { name : string } }) => c.function.name,
);
console. log ( "served by:" , data.model); // the concrete model behind the slug you sent
console. log ( "called :" , called);
if ( ! called. includes ( "escalate_to_human" )) {
throw new Error ( "hard invariant broken: refund above the limit" );
}
Run the suite when something changes
The mechanics are simple once the cases and contracts exist. A few details decide whether the run catches anything.
Trigger the run on the change. Re-run the full case set whenever a prompt, model, tool definition, or retrieval setting changes. A suite that runs only when someone remembers to run it will eventually miss the change that mattered.
Our Ori Eval documentation adds a caution. An eval sends requests to real models and costs money, so put your evals in a separate job, let a person start the job or run it on a schedule, and don’t put it in your normal unit-test job. Both points hold at once. You can measure the cost before you commit to it. ori eval --pilot 1 runs one sampled case per eval file that wraps its case list in pilotCases() and reports the measured cost per model, split between the agent and the judge, alongside an estimate for the full suite. The job you want is scoped to the paths that hold your prompts, model configuration, tool definitions, and retrieval settings, so it triggers on the changes this guide is about and stays quiet for the rest. A failed eval returns a non-zero exit code and fails the job, so a worse agent can stop a release once you make that job a dependency of it. The workflow below goes in your repository at .github/workflows/agent-evals.yml . It pins one Ori release and checks the downloaded binary against a SHA-256 digest written into the workflow, so the job that holds your OPENROUTER_API_KEY runs only the binary you reviewed, and a replaced release asset fails the check. Pick the tag from the Ori releases page , download ori-linux-x64 once, and record its sha256sum output as ORI_SHA256 . The SHA256SUMS file on the same release lists the digest of every asset, and it should match the digest you computed. Our docs also show the one-line installer, curl -fsSL https://openrouter.ai/labs/ori/install.sh | bash , which installs the newest stable release and is the shorter option on a developer machine.
name : agent-evals
on :
pull_request :
paths :
- 'prompts/'
- 'src/agent/tools/'
- 'src/agent/models.ts'
- 'src/retrieval/**'
jobs :
eval :
runs-on : ubuntu-latest
steps :
- uses : actions/checkout@v6
- uses : oven-sh/setup-bun@v2
- name : Install Ori
env :
ORI_RELEASE : cli-0.15.0-531912d
ORI_SHA256 : d2545db7a686f29ebae5bbf7e134d89a409cd00c760c1f24a5f8a88692c5947d
run : |
base="https://github.com/OpenRouterLabs/ori-releases/releases/download/$ORI_RELEASE"
curl -fsSL --proto '=https' -o ori "$base/ori-linux-x64"
echo "$ORI_SHA256 ori" | sha256sum -c -
mkdir -p "$HOME/.local/bin"
install -m 0755 ori "$HOME/.local/bin/ori"
echo "$HOME/.local/bin" >> "$GITHUB_PATH"
- name : Run the evals
run : ori eval --report eval-report.md
env :
OPENROUTER_API_KEY : ${{ secrets.OPENROUTER_API_KEY }}
- name : Add the report to the job summary
if : always()
run : cat eval-report.md >> "$GITHUB_STEP_SUMMARY"
Score the delta as well as the pass. A case that a judge scored well last month and scores lower today hasn’t failed, and it’s still a regression worth opening. Treat a meaningful score drop the way you would treat a failing test. Check that your rubric can move before you rely on it, because a rubric that scores every answer alike reports a clean pass while measuring nothing.
Match the check to the case. Deterministic cases, where you can name the exact tool and argument you expect, get exact or structural checks. Open-ended cases, such as whether an explanation is accurate and correctly scoped, need a judge model, because no single correct string exists to match against.
Ori Eval covers both shapes in one file. Assertions such as run.tool('lookup_order').toBeCalled() , run.toComplete() , run.toCostAtMost(0.01) , and run.toFinishWithin(30_000) handle the structural side. setupJudge({ minScore: 0.8 }) scores open-ended cases from 0 to 1 against criteria you write. Ori also resolves one harness and one model per run and holds them for every test in that run, so two runs of the same eval files use the same configuration.
Test across a model swap
Switching models on OpenRouter is a configuration change rather than a rewrite. That only helps if you can show that behavior stayed put when you made the switch.
Price is usually what starts the conversation. Two models we serve today sit at opposite ends of the price range, and both list tools in their supported parameters.
Model Slug Input per M tokens Output per M tokens Providers
Claude Fable 5.1 anthropic/claude-fable-5.1 $10.00 $50.00 4
Gemini 3.8 Flash google/gemini-3.8-flash $0.75 $3.75 2
Checked September 18, 2026, against the live Claude Fable 5.1 and Gemini 3.8 Flash endpoint data. Gemini 3.8 Flash prices are for the standard tier. Both Google providers also serve flex and priority tiers at different prices. Prices change, so recheck before you plan around a ratio.
A thirteenfold difference in input price is reason enough to try the swap. The run is what earns the right to ship it. The mechanics are the locked case set and the contract you already have, with one variable moved. Pin the prompt, the tool definitions, the tool results, the case set, the judge, and the inference parameters, then run the suite against the candidate before any real traffic reaches it. When a result moves, you know the model moved it.
Pinning includes the slug itself. An alias like ~anthropic/claude-fable-latest routes to the newest concrete model in that family and updates whenever the author publishes a new version. Name the exact version on both sides of the comparison, and read the response’s model field to confirm what served each call.
Pinning also includes the inference parameters, and the two models don’t accept the same ones. Each entry in the models endpoint has a supported_parameters array. Gemini 3.8 Flash lists temperature . Claude Fable 5.1 doesn’t, so with default routing a temperature value sent to it is ignored by the provider rather than applied, and settin
首次收录 · 2026-10-01 · 9.94 分