你更新了一个提示词,或者提供商在同一个模型 ID 下推出了一个新的检查点,导致生产环境中的某些表现出现倒退。上周的用户投诉情况与上周相比略有不同。像 MMLU 这样的公开基准测试无法捕捉到这一点。它衡量的是跨学术科目的通用能力,而不是模型如何处理你产品的流量。
一个黄金评估数据集填补了这一空白。它是一套精心策划的生产输入数据集合,配有经过审核的预期输出结果,通过 Git 进行版本控制,并在每次部署前运行。它回答了一个基准测试无法回答的问题:这次更改是有助于还是损害了你服务的流量?
本指南介绍了什么是黄金集,为什么生产数据比合成数据是更好的基础,以及如何从实时流量中构建黄金集的五个步骤。它还介绍了如何通过一个 API 对许多候选模型运行同一组数据,从而让你能够根据自己流量的证据来选择下一个模型,而不是依据排行榜排名。
简而言之
黄金评估数据集是一套精心策划的生产输入数据集合,配有经过审核的预期输出结果,用作每次重大更改前的回归测试。
一旦该集合存在,你就可以通过一个 API 在多个模型上运行它,并根据你自己流量的证据来选择下一个模型。
规模要与任务相匹配。大约 10 个项目用于探索单个问题,100 到 1,000 个用于完整的回归集。合适的规模取决于指标、方差以及你需要检测的最小差异。
对故障模式的覆盖比数量更重要。生产示例包含了你的用户遇到的故障。合成示例包含了有人想象到的故障。
将数据集、评分标准和基线一起进行版本控制。否则,你无法判断回归是由模型更改、提示词更改还是数据集更改引起的。
什么是黄金评估数据集
黄金评估数据集是一组带有经过审核的预期输出结果的生产示例,这些数据在训练时被隔离出来,并在每个候选发布版本上进行评估。拥有领域知识的人确认了每个预期输出结果,或者决定该项目不需要预期输出结果,因为像格式、安全性或语气这样的无参考检查是通过标准。
这种审核使得该集合变得有用。如果没有它,当分数下降时,你无法判断是更改导致质量倒退,还是第 23 项的数据标签有误。
将其视为行为而非代码的回归测试套件。它在 CI(持续集成)中运行,拥有已知正确的预期输出结果,并捕捉部署之间的漂移。与单元测试的区别在于,预期输出结果是一种判断,因此测试框架必须做的不仅仅是比较字符串。
黄金集不是基准测试、训练集或 A/B 测试。基准测试衡量通用能力。黄金集衡量在你流量上的能力。训练集用于教导模型。黄金集在模型出现倒退时捕捉到它。A/B 测试衡量实时用户结果。黄金集衡量更改是否足够安全,以至于可以对其进行 A/B 测试。
为什么生产数据是更好的基础
合成评估测试的是模型回答有人想象用户可能会问的问题的能力。你需要测试的是你的用户实际提出的问题,以及他们使用的措辞。生产流量是更好的基础,原因有三点。
分布是正确的。如果你的 70% 流量与定价有关,那么你的评估集中也应有 70% 与定价有关。从想象的边缘情况中抽取的合成集无法保留这种分布。保留观察到的分布使得回归分数能够反映对用户的影响。
故障模式正是你关心的那些。用户会发现你想不到去编写的故障模式。一条混合了三种语言的消息、一个粘贴了整个错误日志的支持工单、一个引用了你上个季度重命名的产品功能的查询。除非有人想到了这些,否则合成集中不会包含这些内容。
这些示例在追踪你的产品。六个月前上线、原本用于处理密码重置的支持机器人,现在要应对多账户问题、退款升级投诉以及竞争对手用户界面的截图。从生产环境中提取的“黄金数据集”能够追踪这种漂移现象,而发布时冻结的合成数据集则无法做到。
合成示例仍有其用武之地,但应作为扩展而非基础。利用它们来填补已知故障模式中真实示例不足覆盖的空白。你可以通过调用模型并使用结构化输出来生成这些示例,使每个示例都能落入你的数据集架构中;或者通过对现有项目进行改写。在元数据中标记它们是合成的,并使其在数据集中占少数。
五步流程
第一步:抽取生产流量样本
从一周或两周的日志输入和输出开始。如果该功能具有季节性或使用频率较低,请使用更长的时间窗口。
先进行随机抽样,然后观察结果。如果你的流量具有长尾特征,随机抽样会过度代表头部(高频)部分。这通常是你所期望的,因为头部的回归问题会影响最多的用户,但请检查罕见且关键的意图是否出现。
记录你后续可能需要切片分析的所有字段。这包括意图、功能区域、用户细分、时间戳,以及生成你所采样响应所使用的模型和提示词版本。你无法在事后重建这些元数据。
抽取生产流量意味着抽取用户数据。请检查你的服务条款,在数据进入评估框架之前运行个人身份信息(PII)清理管道,并记录哪些示例经过了转换,以便日后重现清理过程。
目标是在进入第二步时拥有一个包含数百个示例的原始池。你将对其进行大幅精简。
第二步:去重与聚类
生产流量具有重复性。一个支持聊天机器人每天可能会以略有不同的措辞看到上百次相同的“如何重置我的密码”问题。你只需要在数据集中保留其中一个,而不是五十个。
精确匹配的去重会遗漏改写后的表达。为了捕捉这些情况,请检查数据集是否已经覆盖了相同的意图,可以通过对标准化输入进行精确匹配,或通过比较输入之间的嵌入相似度来实现。
一旦获得去重后的池,便在其中抽样以确保覆盖率。如果你的流量中有一半来自某个特定意图,那么该意图应占黄金数据集的大约一半,但不能全部是它。一个代表性不足且表现糟糕的意图,其代价远高于一个代表充分且优雅失败的意图,因此权重应向那些你无法承受的故障模式倾斜。
关于目标规模,Langfuse 的建议是一个有用的起点。将这些数字视为参考点,并根据你的流量和评估目标进行调整。
| 目的 | 典型规模 |
|---|---|
| 探索单一问题 | 约 10 个条目 |
| 测试模型能力边界 | 约 10 个复杂的未解决问题 |
| 较大变更的 CI 检查 | 100 到 1,000 个覆盖生产分布的条目 |
| 使用对抗性提示词的护栏测试 | 规模较大且不断增长,随新案例出现而增加 |
保留一个包含数十到数百个示例的子集用于拉取请求(Pull Request)门禁,以保持其运行速度,并将完整数据集保留给发布分支或夜间构建运行。
第三步:添加预期输出
对于每个输入,必须由合格人员写下正确的输出应该是什么样子。有时这只是一个字符串。更多时候它是一个评分标准(rubric)。哪些事实必须存在,哪些主张不得出现,以及需要什么样的语气或格式。
采用两种评分标准选择可以使评分更加一致。使用二元标准,其中每个标准要么是“满足”(MET),要么是“未满足”(UNMET)。使用分析性评分标准,分别对每个标准进行评分,而不是分配一个总体分数的整体性评分标准。分析性评分标准能告诉你哪些方面变差了,而不仅仅是某件事是否发生了变化。
让两个人独立地对子集进行标注。在出现分歧的地方,将评分标准视为问题所在,而非标注者。修正评分标准,直到两位评审员对同一模型输出给出相同的评分为止。
并非每个项目都需要硬编码的预期输出。仅由无参考评估器(如格式有效性、安全性或语气)检查的项目根本不需要预期输出。黄金集可以混合这两种类型。
步骤 4:运行首次评估并修正评分标准
在你信任该数据集之前,先用当前的生产模型对其进行评估,并仔细查看每一次失败案例。
将失败案例分为三类。真实失败,即模型回答错误且预期输出正确。保留这些。评分标准问题,即模型给出了合理的答案,但你的预期输出未能预见这种情况。修正预期输出。模糊示例,即任何合理的评分标准都无法评判的案例。将其剔除。
预计在此次评估中会剔除部分数据集。具体数量取决于你在步骤 3 中编写预期输出的仔细程度。剔除正是此次评估的目的。一个无论模型如何变化都产生相同通过率的金集并没有衡量任何东西。
不要跳过这一步。如果你在此处未能发现的评分标准问题,会在实际部署时表现为虚假的回归错误。
步骤 5:提交到 Git,接入 CI,迭代
黄金集应属于你的代码库,与提示词和运行它的代码一起进行版本控制。这正是将其从临时质量检查转变为回归测试的关键。
最小的目录结构如下所示。
/evals
/golden
dataset.jsonl # 每行一个示例
rubric.md # 评分方法
run.ts # 加载器和框架
baseline.json # 当前生产模型的通过率
/synthetic
dataset.jsonl # 罕见失败模式,标记为 synthetic: true
将每个示例存储为 JSON 行,包含输入、预期输出、标签和元数据字段。将评分标准存储为人类可读的文档,供评分器读取,以便任何查看拉取请求差异的人都能看出是评分标准的变更导致了通过率的变动。
dataset.jsonl 中的一行如下所示。
{ "id" : "pw-reset-locked-account" , "input" : "i cant log in, tried resetting three times and its saying account locked. help" , "expected" : "The bot should acknowledge the lockout, ask for the account email, and route to the account-recovery flow. It must not offer to reset the password directly." , "rubric" : "must_ask_email;must_route_recovery;must_not_reset_directly" , "tags" :[ "password_reset" , "edge_case" ], "synthetic" : false }
这些字段服务于不同的读者。id 为每个项目提供稳定的引用,使其在 Git 差异中得以保留。input 是生产环境的输入,经过清理后保持原样。expected 是人类可读的文本,LLM 评估器可以阅读。rubric 包含评估器依据进行评分的可机器检查的标准。tags 让你能够按意图切片通过率。synthetic 在聚合指标中将真实项目和合成项目分开。
将框架接入 CI,以便在每次提示词变更或模型交换时运行。如果通过率下降超过定义的阈值,则构建失败;对于较小的下降,则需要书面确认。一个在每个拉取请求上调用你框架的 GitHub Actions 工作流足以开始使用。
定期刷新数据集。每三个月对六个月以上的项目进行回顾,并将其与当前的产品行为进行核对,这是一个合理的默认设置。如果产品变化迅速,则每月回顾一次。过时的黄金集最终将不再能预测生产环境的行为。
与数据一起对评分标准进行版本控制。无法诊断的失败运行毫无用处,而评分标准的变更是事后最难察觉的一类变更。如果有人为了修复看似虚假的回归而编辑了某个标准,该编辑应作为可见的变更出现在拉取请求差异中,而不是在六周后表现为无法解释的通过率变动。
不要因为项目持续通过就将其退役。一个通过的项目证明了该行为仍然有效。仅当该项目测试的行为在产品中不再存在,或预期输出现在错误时,才退役该项目。
通过单一 API 进行跨模型基准测试
当你拥有包含预期输出的 100 个示例后,将同一组数据针对不同的模型运行,就能告诉你该模型在你的工作负载上的表现。这不是在 MMLU 上测试,也不是在别人的代码评估框架中测试,而是在你的用户实际发送的问题上进行测试。
跨供应商比较模型通常意味着每个候选模型都有各自的 SDK、认证方式、响应格式和速率限制。通过 OpenRouter,相同的 OpenAI 兼容请求体可以在整个模型目录中通用。对于共享接口并支持你请求中使用的参数的候选模型,你只需将 model: "openai/gpt-5.1" 更改为 model: "anthropic/claude-fable-5.1" 并重新运行评估框架即可进行比较。如果模型在上下文长度、工具支持或支持的参数方面存在差异,则需要针对每个候选模型进行一些集成工作。models 端点的每个条目都列出了模型的上下文长度和 supported_parameters 数组,因此你可以在切换前进行检查。
一个最小的评估框架如下所示。
import OpenAI from "openai" ;
import { readFileSync } from "fs" ;
const client = new OpenAI ({
apiKey: process.env. OPENROUTER_API_KEY ?? "" ,
baseURL: "https://openrouter.ai/api/v1" ,
});
const dataset = readFileSync ( "./evals/golden/dataset.jsonl" , "utf-8" )
. trim ()
. split ( " \n " )
. map (( line ) => JSON . parse (line));
const modelsToTest = [
"openai/gpt-5.1" ,
"anthropic/claude-fable-5.1" ,
"google/gemini-3.8-flash" ,
];
const results : Record < string , { pass : number ; total : number }> = {};
for ( const model of modelsToTest) {
results[model] = { pass: 0 , total: 0 };
for ( const example of dataset) {
const response = await client.chat.completions. create ({
model,
messages: [{ role: "user" , content: example.input }],
});
const output = response.choices[ 0 ].message.content ?? "" ;
const passed = grade (output, example.expected, example.rubric);
results[model].total += 1 ;
if (passed) results[model].pass += 1 ;
}
}
console. table (results);
grade 函数是存放你评分标准的地方。对于字符串匹配的情况,它可以是 output.includes(expected)。对于基于标准的评分,通常是向一个强大的评判模型发起另一次 LLM 调用,传入评分标准和响应,询问该响应是否满足每个标准。
发送给评判模型的评分标准检查提示词如下所示。
评分标准(每项为 MET 或 UNMET):
{criteria}
待评分的响应:
{response}
对于每个标准,在单独的一行输出 MET 或 UNMET。不要包含任何正文。
评判模型会为每个标准返回一个 MET 或 UNMET,grade() 函数会将其解析为:如果所有标准都满足则通过,否则失败。二元标准和“无正文”指令确保了评判输出的可解析性,使其在多次运行中保持一致。
对于设计良好的、具有二元标准和锚点示例的评分标准,LLM 评判模型在检测回归问题方面足够可靠。与人类评分员的一致性因任务、评分标准和评判模型而异,因此在依赖其分数之前,应使用一组人工标记的示例对评判模型进行校准。仅凭随机种子不同,评判模型的可靠性也会发生变化。一项研究让三个 LLM 评判模型对 BIG-Bench Hard 问题的响应各评分 100 次,除了改变随机种子外什么都不变,并测量了评判模型之间的评分者间可靠性,在不同复现中范围从 0.167 到 1.00 不等。在该量表上,1.0 表示完美一致。使用与待测模型不同的模型作为评判模型,以避免自我偏好偏差,并且不要将单次评判运行视为高风险决策的事实真相。
在每个模型上对同一组数据进行评分。如果在运行之间更改数据集,你就是在同时测量两件事。冻结该数据集,记录提交哈希值,然后交换模型。
既要关注成本,也要关注质量。在你的黄金数据集中最准确的模型,每次调用的成本可能是第二准确模型的 30 倍,而质量差异仅为两个百分点。将这种权衡关系呈现给负责预算的人。在我们的模型比较页面上,你可以在运行完整基准测试之前查看两个模型之间的价格差距。
在需要可复现性时,请使用提供商路由控制。相同的模型标识符可以路由到不同的提供商,而不同提供商的行为可能存在细微差异。要将基准测试运行固定到某一特定提供商,请在 pr 中设置 order 字段。
You update a prompt, or a provider rolls out a new checkpoint under the same model ID, and something in production regresses. The last week of user complaints looks slightly different from the week before. A public benchmark like MMLU won’t catch that. It measures general capability across academic subjects, not how the model handles your product’s traffic.
A golden eval dataset closes that gap. It’s a curated collection of production inputs paired with reviewed expected outputs, versioned in Git, and run before every deploy. It answers the question a benchmark can’t. Does this change help or hurt on the traffic you serve?
This guide covers what a golden set is, why production data is a better foundation than synthetic data, and a five-step process for building one from live traffic. It also covers how to run the same set against many candidate models through one API, so you can pick your next model on evidence from your own traffic rather than leaderboard rank.
Tl;dr
A golden eval dataset is a curated set of production inputs paired with reviewed expected outputs, used as a regression test before every meaningful change.
Once the set exists, you can run it across many models through one API and pick your next model on evidence from your own traffic.
Match the size to the job. Roughly 10 items to explore a single issue, 100 to 1,000 for a full regression set. The right size depends on the metric, the variance, and the smallest difference you need to detect.
Coverage of failure modes matters more than volume. Production examples carry the failures your users hit. Synthetic examples carry the failures someone imagined.
Version the dataset, the rubric, and the baseline together. Otherwise you can’t tell whether a regression came from a model change, a prompt change, or a dataset change.
What a golden eval dataset is
A golden eval dataset is a collection of production examples with reviewed expected outputs, held out from training and evaluated on every release candidate. Someone with domain knowledge confirmed each expected output, or decided the item doesn’t need one because a reference-free check like format, safety, or tone is the pass criterion.
That review is what makes the set useful. Without it, when a score drop appears, you can’t tell whether the change regressed quality or item 23 has a wrong label.
Treat it as a regression test suite for behavior rather than code. It runs in CI, it has known-good expected outputs, and it catches drift between deploys. The difference from a unit test is that the expected output is a judgment, so the harness has to do more than compare strings.
A golden set is not a benchmark, a training set, or an A/B test. A benchmark measures general capability. A golden set measures capability on your traffic. A training set teaches the model. A golden set catches the model when it regresses. An A/B test measures live user outcomes. A golden set measures whether a change is safe enough to run an A/B test on at all.
Why production data is the better foundation
Synthetic evals test a model’s ability to answer questions someone imagined a user might ask. What you need to test is the questions your users asked, in the phrasing they used. Production traffic is a better foundation for three reasons.
The distribution is right. If 70 percent of your traffic is about pricing, 70 percent of your eval set should be about pricing. A synthetic set drawn from imagined edge cases doesn’t preserve that distribution. Preserving the observed distribution is what makes a regression score reflect user impact.
The failure modes are the ones you care about. Users find failure modes you wouldn’t think to write. A message that mixes three languages, a support ticket that pastes an entire error log, a query that references a product feature you renamed last quarter. A synthetic set won’t contain these unless someone thought of them.
The examples track your product. A support bot that launched handling password resets six months ago now gets multi-account questions, refund escalations, and screenshots of a competitor’s UI. A golden set drawn from production tracks that drift. A synthetic set frozen at launch doesn’t.
Synthetic examples still have a place as an extension rather than a foundation. Use them to fill coverage gaps for known failure modes you have too few real examples of. You can generate them with a model call that uses structured outputs so each example lands in your dataset schema, or by paraphrasing existing items. Mark them as synthetic in metadata, and keep them a minority of the set.
The five-step process
Step 1: Pull a sample of production traffic
Start with a week or two of logged inputs and outputs. Use a longer window if the feature is seasonal or thinly used.
Sample randomly first, then look at what came out. If your traffic has a long tail, a random sample over-represents the head. That’s often what you want, since regressions in the head affect the most users, but check whether rare and critical intents show up at all.
Log every field you might want to slice on later. That includes the intent, the feature area, the user segment, the timestamp, and the model and prompt version that produced the response you sampled. You can’t reconstruct this metadata later.
Sampling production traffic means sampling user data. Check your terms of service, run your PII scrubbing pipeline before the data reaches an eval harness, and log which examples you transformed so you can reproduce the scrubbing later.
Aim for a raw pool of a few hundred examples going into step 2. You’ll trim heavily.
Step 2: Deduplicate and cluster
Production traffic is repetitive. A support chatbot might see the same “how do I reset my password” question a hundred times a day in slightly different phrasings. You want one of those in the set, not fifty.
Exact-match deduplication misses paraphrases. To catch them, check whether the dataset already covers the same intent, either by exact-matching normalized inputs or by comparing embedding similarity between inputs.
Once you have a deduplicated pool, sample within it for coverage. If half your traffic is one intent, that intent should get roughly half the golden set, but not all of it. An under-represented intent that fails badly costs more than a well-represented intent that fails gracefully, so weight toward the failure modes you can’t afford.
For target size, Langfuse’s guidance is a useful starting point. Treat the numbers as reference points and adjust to your traffic and evaluation goal.
Purpose Typical size
Exploring a single issue About 10 items
Testing model capability boundaries About 10 complex unsolved examples
CI checks on larger changes 100 to 1,000 items covering the production distribution
Guardrail testing with adversarial prompts Large and growing, add cases as they surface
Keep a subset of tens to low hundreds for the pull request gate so it stays fast, and reserve the full set for release branches or nightly runs.
Step 3: Add expected outputs
For every input, someone qualified has to write down what the right output looks like. Sometimes that’s a single string. More often it’s a rubric. Which facts must be present, which claims must not be made, and what tone or format is required.
Two rubric choices make grading more consistent. Use binary criteria, where each criterion is either MET or UNMET. Use analytic rubrics, which score each criterion separately, instead of holistic rubrics that assign one overall score. Analytic rubrics tell you what got worse, not only whether something did.
Have two people annotate a subset independently. Where they disagree, treat the rubric as the problem, not the annotators. Fix the rubric until two reviewers grading the same model output would grade it the same way.
Not every item needs a hard-coded expected output. Items checked only by reference-free evaluators, such as format validity, safety, or tone, need no expected output at all. A golden set can mix both kinds.
Step 4: Run a first evaluation and fix the rubric
Before you trust the set, run your current production model against it and look at every failure.
Sort the failures into three categories. Real failures, where the model got it wrong and the expected output is right. Keep those. Rubric problems, where the model gave a reasonable answer that your expected output didn’t anticipate. Fix the expected output. Ambiguous examples that no reasonable rubric can grade. Cut them.
Expect to cut part of the set on this pass. How much depends on how carefully you wrote the expected outputs in step 3. The cut is the point of the pass. A golden set that produces the same pass rate regardless of the model change isn’t measuring anything.
Don’t skip this step. A rubric problem you don’t catch here shows up later as a false regression during a real deploy.
Step 5: Commit to Git, wire into CI, iterate
The golden set belongs in your codebase, versioned alongside the prompts and the code that runs it. That’s what turns it from an ad hoc quality check into a regression test.
A minimal directory layout looks like this.
/evals
/golden
dataset.jsonl # one example per line
rubric.md # how to grade
run.ts # loader and harness
baseline.json # pass rates on the current production model
/synthetic
dataset.jsonl # rare failure modes, marked with synthetic: true
Store each example as a JSON line with fields for the input, the expected output, tags, and metadata. Store the rubric as a human-readable document that the grader reads from, so anyone reading a pull request diff can see whether a rubric change is what moved the pass rate.
A single line in dataset.jsonl looks like this.
{ "id" : "pw-reset-locked-account" , "input" : "i cant log in, tried resetting three times and its saying account locked. help" , "expected" : "The bot should acknowledge the lockout, ask for the account email, and route to the account-recovery flow. It must not offer to reset the password directly." , "rubric" : "must_ask_email;must_route_recovery;must_not_reset_directly" , "tags" :[ "password_reset" , "edge_case" ], "synthetic" : false }
The fields serve different readers. id gives each item a stable reference that survives Git diffs. input is the production input, verbatim after scrubbing. expected is human-readable prose that an LLM judge can read. rubric holds the machine-checkable criteria the judge grades against. tags let you slice pass rates by intent. synthetic keeps real and synthetic items separate in aggregate metrics.
Wire the harness into CI to run on every prompt change or model swap. Fail the build if the pass rate drops more than a defined threshold, and require a written acknowledgment on any smaller drop. A GitHub Actions workflow that calls your harness on every pull request is enough to start.
Refresh the set on a cadence. A quarterly review of items older than six months, checking each against current product behavior, is a reasonable default. Review monthly if the product is changing fast. A stale golden set eventually stops predicting production behavior.
Version the rubric alongside the data. A failing run is only useful if you can diagnose it, and rubric changes are the hardest kind of change to catch after the fact. If someone edits a criterion to fix what looked like a false regression, that edit should appear as a visible change in a pull request diff, not as an unexplained pass rate move six weeks later.
Don’t retire items because they keep passing. A passing item is proving that the behavior still holds. Retire an item only when the behavior it tests no longer exists in the product, or when the expected output is now wrong.
Cross-model benchmarking through one API
Once you have 100 examples with expected outputs, running the same set against a different model tells you how that model performs on your workload. Not on MMLU, and not on someone else’s coding harness. On the questions your users send.
Comparing models across vendors usually means a different SDK, different authentication, a different response shape, and different rate limits for each candidate. Through OpenRouter , the same OpenAI-compatible request body works across the model catalog . For candidates that share an interface and support the parameters your request uses, you can compare them by changing model: "openai/gpt-5.1" to model: "anthropic/claude-fable-5.1" and rerunning the harness. Where models differ on context length, tool support, or supported parameters, expect some integration work per candidate. Each entry in the models endpoint lists the model’s context length and a supported_parameters array, so you can check before switching.
A minimal harness looks like this.
import OpenAI from "openai" ;
import { readFileSync } from "fs" ;
const client = new OpenAI ({
apiKey: process.env. OPENROUTER_API_KEY ?? "" ,
baseURL: "https://openrouter.ai/api/v1" ,
});
const dataset = readFileSync ( "./evals/golden/dataset.jsonl" , "utf-8" )
. trim ()
. split ( " \n " )
. map (( line ) => JSON . parse (line));
const modelsToTest = [
"openai/gpt-5.1" ,
"anthropic/claude-fable-5.1" ,
"google/gemini-3.8-flash" ,
];
const results : Record < string , { pass : number ; total : number }> = {};
for ( const model of modelsToTest) {
results[model] = { pass: 0 , total: 0 };
for ( const example of dataset) {
const response = await client.chat.completions. create ({
model,
messages: [{ role: "user" , content: example.input }],
});
const output = response.choices[ 0 ].message.content ?? "" ;
const passed = grade (output, example.expected, example.rubric);
results[model].total += 1 ;
if (passed) results[model].pass += 1 ;
}
}
console. table (results);
The grade function is where your rubric lives. For string-match cases, it can be output.includes(expected) . For rubric-based grading, it’s usually another LLM call to a strong judge model with the rubric and the response, asking whether the response satisfies each criterion.
A rubric-check prompt to the judge looks like this.
Rubric criteria (each MET or UNMET):
{criteria}
Response to grade:
{response}
For each criterion, output MET or UNMET on its own line. No prose.
The judge returns one MET or UNMET per criterion, which grade() parses into a pass if every criterion is MET and a fail otherwise. Binary criteria and the “no prose” instruction keep the judge output consistent enough to parse across runs.
An LLM judge is reliable enough for regression detection on a well-designed rubric with binary criteria and anchor examples. Agreement with human graders varies by task, rubric, and judge model, so calibrate the judge against a set of human-labeled examples before you rely on its scores. Judge reliability also varies with the random seed alone. One study had three LLM judges grade responses to BIG-Bench Hard questions 100 times each, changing nothing but the seed, and measured inter-rater reliability between the judges ranging from 0.167 to 1.00 across replications. On that scale, 1.0 is perfect agreement. Use a different model as the judge than the model under test to avoid self-preference bias, and don’t treat a single judge run as ground truth on high-stakes decisions.
Score the same set on every model. If you change the dataset between runs, you’re measuring two things at once. Freeze the set, record the commit hash, then swap models.
Watch cost as well as quality. The most accurate model on your golden set might cost 30 times as much per call as the second-most-accurate model, for a quality difference of two percentage points. Surface that tradeoff to whoever owns the budget. Our model comparison page shows the price gap between two models before you run the full benchmark.
Use provider routing controls when you need reproducibility. The same model slug can route to different providers, and providers can differ slightly in behavior. To pin a benchmark run to one provider, set the order field in the pr
首次收录 · 2026-10-01 · 9.94 分