Jev 是 TypeSafe 推出的一种决策模型。它接收非结构化输入,并返回带有概率的键入式判断结果,随后由你的代码决定下一步的操作。本指南展示了如何在自己的问题上使用 Jev,并通过一个示例从头到尾进行了完整演示:在 TypeScript 中审核市场平台上的商品列表。
Jev 的工作原理
好的,既然你已经了解了 Jev 的用途,并且感觉正走在正确的道路上。那么,问题在于:Jev 请求在底层看起来是什么样的?
让我们从头开始。每个 Jev 请求都由一个状态(state)和一组问题(questions)组成。状态是 Jev 应该评估的输入。它可以是一个字符串、JSON 对象或字符串数组。如果你试图评估一系列项目,例如对话中的消息,那么使用数组将是最佳选择。问题是 Jev 应对状态执行的一系列判断。一个问题由两个元素组成,即指令(instructions)和标准(criteria)。指令描述了应执行的判断。标准描述了可能的答案。在评估状态时,Jev 将确定每个标准与状态的契合程度,并将结果返回。当状态为 JSON 对象时,指令可以通过用反引号包裹来引用状态中的命名字段。我们将这称为字段路径(field path)。例如,有效的字段路径可以是 listing.description。当提供此指令时,Jev 将读取命名字段中的值并对其进行判断。
问题类型包括 choice、noul 和 score。choice 问题从你定义的一组选项中选择一个。noul 问题决定一个陈述是否为真。score 问题将主体放置在你定义的有序等级量表上。
Jev 针对相同的状态单独且并行地评估你的每个问题。它只会为每个问题提供一个答案。每个答案都有一个类型,即 choice、noul 或 score 之一。choice 答案包括:1. 获胜选项;2. 每个选项的概率;3. TypeSafe 根据概率分布的形状计算出的介于 0 到 1 之间的置信度(confidence)。如果概率高度集中在获胜选项上,置信度将很高。如果概率分散开来,置信度将会较低,即使获胜选项单独的概率最高。置信度并非获胜选项的概率,而是对整个分布的总结。noul 答案包括:1. 陈述为真的单一概率;2. 没有置信度,因为 noul 概率本身已经表达了模型的自信程度。score 答案包括:1. 等级数字的概率加权平均值;2. 每个等级数字的概率;3. 一个置信度;4. 将等级数字映射到相应描述的图例。
每个答案都是一个选项或一个数字,你的代码可以对其进行比较、设定阈值和组合。不会为你生成需要解析的文本。这就是为什么比起要求聊天模型回答“是”或“否”,更倾向于使用 Jev 的原因。你得到的是对你所定义答案的概率分布,因此不确定性是一个你可以用于路由的数字。
Jev 在 OpenRouter Decisions API 上运行,使用 typesafe/jev-1.13 模型 ID。任何对该 API 的使用都将计入你的 OpenRouter 账户费用。对 Jev 的请求仅由文本组成,令牌预算同时消耗于状态和所提出的问题。Jev 模型页面目前显示了预算和定价。新用户可以按照 Jev 教程在几分钟内获得第一个答案。
以市场平台作为运行示例
假设你正在构建一个让用户可以发布二手物品出售信息的交易市场。卖家需要撰写标题和描述,选择类别,注明物品状况,并上传照片。在列表发布之前,你需要确保没有人列出违禁物品,或要求通过平台外进行支付。同时,你也希望捕捉到归类错误的物品,以及标题或声明状况与描述相矛盾的情况。
其中一些检查只是普通的代码逻辑,但另一些则需要实际阅读列表内容并解读其含义。这正是 Jev 的用途所在。下文每一节描述了在一般情况下使用 Jev 的一个部分,然后说明它在交易市场中的具体工作方式。相同的模式同样适用于工单路由、代理工具门控或你正在构建的任何其他应用。
Jev 如何融入程序
Jev 是一个 System One 模型。该术语是 TypeSafe 对产生类型化决策和校准概率的模型的命名。TypeSafe 的构建指南描述了如何使用此类模型进行开发。在原本正常的软件工作流链条中,插入 Jev 以处理需要判断的环节。代码负责控制流、确定性规则和副作用。Jev 为代码准备的一组少量问题提供类型化、基于状态的答案。代码将这些答案转化为行动。Jev 不会自行决定其下一步行动。
在市场交易中,这一原则在卖家点击发布和列表上线之间创建了一个管道。它接收一个列表,并返回“发布”、“由人工暂存”或“拒绝并说明理由”。
硬性规则首先以代码形式运行,例如照片数量、价格范围和最小文本长度。未能通过硬性规则的列表会在发送给 Jev 之前被直接拒绝。
状态构建将卖家的文本与代码计算出的事实相结合,并决定 Jev 所看到的内容。
判断环节通过一次请求向 Jev 提出五个关于该状态的类型化问题。Jev 返回概率,而从不返回行动指令。
验证确保响应符合策略预期的形状,如果不符合则停止处理。
策略将这些概率与基于审核人员已做出决定的列表所校准的阈值进行比较。
第 3 步是 Jev,其余四步由你负责。这对于所有 Jev 集成都是相同的。Jev 是一个从状态到证据的函数,而其余部分都是普通软件。
决定由 Jev 判断的内容以及由你的代码决定的内容
集成 Jev 的第一个设计问题是,应该发送什么给 Jev 进行判断。测试的标准是:代码是否可以在不阅读任何文本的情况下计算出答案。如果可以(例如计数、日期比较、查找),则将其保留在代码中,因为结果是精确且免费的。Jev 用于需要解读自然语言才能进行的判断。
TypeSafe 关于 Jev 1.13 的 jaggedness 页面将算术运算、精确计数和日期比较列为应留在代码中而非放入问题中的项目。
应用于交易市场时,该测试对检查项的分类如下。
| 检查项 | 存放位置 | 原因 |
|---|---|---|
| 价格在范围内,至少有一张照片,达到最小长度 | 代码 | 确定性逻辑。模型的可信度永远低于 if 语句。 |
| 价格远低于该类别通常的售价 | 代码 | 基于自有销售数据的算术运算。将结果作为事实传递给 Jev。 |
| 物品属于违禁类型(武器、假冒商品、召回的儿童座椅) | Jev | 卖家很少使用违禁词。“镜面品质,同厂生产”绝不会直接表明是假冒商品。 |
| 列表要求买家在平台外支付或聊天 | Jev | 电话号码很容易通过正则表达式识别。“我也可以通过电话谈这笔交易”则不行。 |
| 物品符合卖家选择的类别 | Jev | 基于定义的语义判断。 |
| 描述与标题或声明的状况相矛盾 | Jev | 需要比较含义。 |
这些检查项也定义了每个阶段共享的列表类型。每个字段属于以下三类之一:硬性规则输入、为 Jev 计算的事实,或 Jev 阅读的文本。状况按从最差到最好的顺序排列,因此声明状况的索引与后续状况分数的层级相对应。
// listing.ts
export const CATEGORIES = {
electronics: '手机、笔记本电脑、相机、音频设备、游戏主机及其配件。',
furniture: '桌子、椅子、沙发、床、置物架及其他家用家具。',
clothing: '服装、鞋履、包袋及时尚配饰。',
sporting_goods: '自行车、健身器材、露营装备及体育运动器材。',
toys_and_baby: '玩具、游戏、婴儿车、汽车安全座椅、婴儿床及其他儿童用品。',
} as const ;
export type Category = keyof typeof CATEGORIES ;
export const CONDITIONS = [ 'for_parts' , 'fair' , 'good' , 'like_new' , 'new' ] as const ;
export type Condition = ( typeof CONDITIONS )[ number ];
export type Listing = {
id : string ;
title : string ;
description : string ;
category : Category ;
condition : Condition ;
priceUsd : number ;
photoCount : number ;
};
类别定义以散文形式编写,因为 Jev 在判断某个商品列表是否属于该类别时,会从状态中读取这些定义。这是一项通用原则。如果判断依赖于你的某条规则,请以散文形式将该规则写入状态,并在你的问题中引用它,以便模型使用你的定义而非其自身的理解。
硬编码规则是标准代码,在任何对 Jev 的调用之前运行。
// rules.ts
import type { Category, Listing } from './listing' ;
// 各类别的典型售价(美元)。在生产环境中,这些数据来自你自己的销售数据。
const TYPICAL_PRICE_USD : Record < Category , number > = {
electronics: 180 ,
furniture: 120 ,
clothing: 35 ,
sporting_goods: 90 ,
toys_and_baby: 40 ,
};
export type PriceSignal = 'far_below_typical' | 'typical' | 'far_above_typical' ;
export function priceSignal ( listing : Listing ) : PriceSignal {
const ratio = listing.priceUsd / TYPICAL_PRICE_USD [listing.category];
if (ratio < 0.2 ) return 'far_below_typical' ;
if (ratio > 10 ) return 'far_above_typical' ;
return 'typical' ;
}
// 硬编码规则。此处包含的任何内容都是你的代码已知的既定事实,因此不涉及模型。
export function hardRuleFailures ( listing : Listing ) : string [] {
const failures : string [] = [];
if (listing.photoCount < 1 ) failures. push ( '至少需要一张照片' );
if (listing.priceUsd < 0.01 || listing.priceUsd > 50_000 ) failures. push ( '价格必须在 $0.01 到 $50,000 之间' );
if (listing.title. trim (). length < 8 ) failures. push ( '标题至少需要 8 个字符' );
if (listing.description. trim (). length < 30 ) failures. push ( '描述至少需要 30 个字符' );
return failures;
}
价格界限在一个检查中展示了划分。这些界限由代码实现,但如果一个通常售价为 $1,800 的类别中出现标价 $120 的手提包,则指向假冒嫌疑,而解读这一信号属于判断范畴。代码执行比较并将最终标签发送给 Jev。
结构化 Jev 所看到的状态
状态是代码与 Jev 之间接口的输入端。TypeSafe 的状态文档将其描述为专家组在接到判断请求前会收到的信息。除了通过状态之外,Jev 对主题一无所知,且请求中的每个问题都会看到相同的状态。
构建状态需遵循适用于任何领域的三条规则。为你的字段使用描述性名称,因为问题通过路径引用字段,而这些名称向模型传达含义。仅包含问题所需的信息,因为无关的上下文会分散模型的注意力。
将计算得出的事实作为已完成的标签传递。Jev 可以直接使用 price_signal: 'far_below_typical',而如果传入原始值 price_usd: 120 和 typical_price: 1800,则等于让它解决一道除法题。
以下是一个市场商品列表的状态示例。
{
"listing" : {
"title" : "Louis Vuitton Neverfull MM 手提包" ,
"description" : "镜面品质 1:1,与专柜版同厂生产。无人能分辨差异。请通过 WhatsApp 联系我获取更多照片和更优惠的价格。" ,
"category" : "服装" ,
"declared_condition" : "全新" ,
"price_usd" : 120
},
"category_definition" : "服装、鞋履、包袋及时尚配饰。" ,
"price_signal" : "典型"
}
卖家 ID、时间戳和照片 URL 被省略,因为这五个问题均未使用它们。包含 category_definition 是因为“类别匹配”问题将其与商品列表进行比较,而 price_signal 是因为“假冒”标准会读取它。
为每个判断选择问题类型
每个 Jev 问题都有一个类型,且该类型应由答案的含义决定。它固定了答案的形态,而形态决定了你后续能对其做什么。
当需要从你定义的一组选项中选出获胜者,且你的代码需要知道是哪一个是时,请使用 choice(选择题)。你可以获得每个选项的获胜结果及其概率,从而观察有多少概率落在了第二名及之后的选项上。当答案是一个非真即假的命题,且你只需要一个用于设定阈值的单一概率时,请使用 noul(是非题)。
你可以使用 score(评分题)来表示程度。它位于一个有序量表上。结果是一个数字,你可以用它与其他数字相减、比较,从而衡量它们之间的差异程度。
不同问题类型的标准格式有所不同。choice 问题接收一个对象,将每个选项名称映射到其描述,最多支持 255 个选项。noul 问题接收一个可选对象,包含 true 和 false 的描述。理想的问题应简洁明了,因此许多问题甚至不需要描述对象。score 问题接收一个有序的水平描述数组(包含 2 到 10 个水平),答案通过该数组的从零开始的索引来衡量。
市场平台使用了这三种类型。“违禁物品”是一个在五类命名类别和一个“无”选项之间的选择,因为政策必须知道违反了哪条规则(武器与假冒商品不同,它们触发的消息和接收审查的人员也不同)。“类别匹配”、“标题与描述矛盾”以及“站外联系”是 noul 问题,因为每个都是是非命题,而政策所需的唯一数字是概率。“描述状态”是一个五个水平的评分,因为政策询问的是声明状态与描述状态之间的差距有多大,而距离只能在有序量表上测量。
独立的问题应放在一个请求中。Jev 针对相同的状态并行评估它们,因此五个问题只需一次往返,而触发其中两个问题的商品列表(本文稍后提到的手提包)将返回两个不同的原因。宽泛的问题在一个答案中隐藏了多个判断。将宽泛问题分解为狭窄、原子化的问题,使其在代码中更容易检查、调整和组合。
编写经得起字面解读的 Jev 标准
指令说明要决定什么,而标准定义答案。Jev 对两者都进行字面应用。这种字面主义是使 Jev 可预测的一部分,这也意味着标准承担了大部分工作。
在为任何领域编写标准时,有几条规则适用。使用输入文本中现有的词汇来描述选项,因为命名了情境的标准(“付款后发送登录详情的共享账户”)才能匹配从未使用你政策术语的文本。对于每个 noul,描述正反两面,以便接近边缘的情况落在正确的一侧。为每个 choice 提供明确的 none 选项,以防万一没有任何情况符合,这样模型就不会被迫选择一个任意项。
对于 score,将每个水平描述为一个具体情境,因为这些水平构成了量表。而像 prohibited 这样的问题 ID 是你的代码读取答案的键,所以不要指望模型会将其解释为指令。
客户端、条件规模和禁止物品标准如下。nou 标准占据请求本身。请在下一节中进一步查找它们。
// judge.ts
import { OpenRouter } from '@openrouter/sdk' ;
import { CATEGORIES, CONDITIONS, type Listing } from './listing' ;
import { priceSignal } from './rules' ;
const openrouter = new OpenRouter ({
apiKey: process.env. OPENROUTER_API_KEY , // 仅限服务器端
});
export const PROHIBITED = {
none: '普通二手或全新物品,一般市场允许出售。' ,
weapon_or_weapon_part:
'枪支、枪支部件或配件、弹药、电击枪,或作为战斗或自卫用途营销的刀具或工具。' ,
medication_or_medical_claim:
'处方药、受控物质,或任何声称具有治疗功效的产品,
Jev is a decision model from TypeSafe. It takes unstructured input and gives you typed judgments with probabilities in return, and your code decides what to do next. This guide shows how to use Jev on a problem of your own, worked end to end on one example, moderating listings on a marketplace in TypeScript.
How Jev works
Okay, so you understand what Jev is for and feel like you’re on the right track. Now, the question is: what does a Jev request look like under the hood?
Let’s start from the beginning. Every Jev request consists of a state and a set of questions . The state is the input that Jev should evaluate. It could be a string, JSON object, or array of strings. If you’re attempting to evaluate a sequence of items, such as messages in a conversation, the array is going to be the way to go. The questions are a list of judgments that Jev should perform on the state. A question consists of two elements, namely, instructions and criteria . The instructions describe the judgment that should be performed. The criteria describe the possible answers. When evaluating the state, Jev will determine how well each of the criteria fits the state and return that as a result. When the state is a JSON object, the instructions can refer to a named field in the state by wrapping it in backticks. We refer to this as the field path. For example, a valid field path could be listing.description . When this instruction is provided, Jev will read the value in the named field and perform the judgment on it.
The question types are choice , noul , and score . A choice question picks one option from a set defined by you. A noul question decides whether a statement is true. A score question places the subject on an ordered scale of levels you define.
Jev evaluates each of your questions individually and in parallel against the same state. It will give you only one answer per question. Each answer has a type , which is either a choice , a noul , or a score . A choice answer consists of: 1. The winning option, 2. The probability of each option, 3. A confidence between 0 and 1 that TypeSafe computes from the shape of the probability distribution. If the probability is highly concentrated in the winning option, confidence will be high. If it is spread out, it will be lower, even if the winning option has the highest probability individually. The confidence is not the probability of the winning option. It’s a summary of the entire distribution. A noul answer consists of: 1. A single probability that the statement is true, 2. No confidence, because the noul probability already expresses how confident the model is. A score answer consists of: 1. The probability-weighted average of the level numbers, 2. The probability of each level number, 3. A confidence, and 4. A legend mapping level numbers to the corresponding descriptions.
Each answer is an option or a number that your code can compare, threshold, and combine. No text is generated for you to parse. This is the reason to prefer Jev over a chat model asked for YES or NO. You get a distribution over the answers you defined, so uncertainty is a number you can route on.
Jev runs on the OpenRouter Decisions API using the typesafe/jev-1.13 model ID. Any usage against the API is billed to your OpenRouter account. A request to Jev consists solely of text, with a token budget being consumed by both the state and the questions asked. The Jev model page currently shows the budget and pricing. New users can follow the Jev tutorial to get their first answer in a few minutes.
A marketplace as the running example
Suppose you’re building a marketplace where users can list used items for sale. Sellers write a title and a description, select a category, note the item’s condition, and upload photos. Before a listing is published you want to make sure nobody is listing a prohibited item or asking to be paid off the platform. You also want to catch items in the wrong category and descriptions that contradict the title or the declared condition.
Some of these checks are just plain code, but others require actually reading the listing and interpreting what it means. This is what Jev is for. Each section below describes one part of using Jev in general, then how it works for the marketplace. The same pattern applies to ticket routing, agent tool gating, or whatever you’re building.
How Jev fits into a program
Jev is a System One model. That term is TypeSafe’s name for a model that produces typed decisions and calibrated probabilities. TypeSafe’s building guide describes how to build with such a model. Insert Jev where judgment is needed in a chain of otherwise normal software workflow. Code handles the control flow, deterministic rules, and side effects. Jev provides typed, state-based answers to a small set of questions prepared by code. Code transforms those answers into actions. Jev doesn’t decide its next action on its own.
In the marketplace, this principle creates a pipeline between a seller clicking publish and the listing going live. It takes a listing and returns publish, hold for a human, or reject with reasons.
Hard rules run first, in code, such as number of photos, price bounds, and minimum text length. Listings that fail them are rejected before anything is sent to Jev.
State construction combines the seller’s text with facts computed by the code, and decides what Jev sees.
Judgment is one request asking Jev five typed questions about that state. Jev returns probabilities and never an action.
Validation ensures the response matches the shape the policy expects, and halts processing if not.
Policy compares those probabilities to thresholds calibrated on listings moderators have already decided.
Stage 3 is Jev and the other four are yours. This is the same for all Jev integrations. Jev is a state-to-evidence function, and everything else is ordinary software.
Decide what Jev decides and what your code decides
The first design question in integrating Jev is what should be sent to it for judgment. The test is whether the code can compute the answer without reading any text. If it can (a count, a date comparison, a lookup), keep it in code, where the result is exact and free. Jev is for judgment that requires interpreting natural language.
TypeSafe’s jaggedness page for Jev 1.13 lists arithmetic, exact counting, and date comparison as items that should stay in code rather than go into questions.
Applied to the marketplace, the test sorts the checks like this.
Check Where it lives Why
Price within bounds, at least one photo, minimum lengths Code Deterministic. A model can only be less reliable than an if .
Price far below what this category usually sells for Code Arithmetic over your own sales data. Pass the result to Jev as a fact.
Item is a prohibited type (weapon, counterfeit, recalled child seat) Jev Sellers rarely use the banned word. “Mirror quality, same factory” never says counterfeit.
Listing asks the buyer to pay or chat off-platform Jev Phone numbers are easy to regex. “I can also do the deal by phone” is not.
Item fits the category the seller chose Jev Semantic judgment against a definition.
Description contradicts the title or declared condition Jev Requires comparing meaning.
The checks also define the listing type every stage shares. Every field is one of three things, a hard-rule input, a fact computed for Jev, or text Jev reads. Conditions are ordered from worst to best, so the index of a declared condition lines up with the levels of the condition score later.
// listing.ts
export const CATEGORIES = {
electronics: 'Phones, laptops, cameras, audio gear, game consoles, and their accessories.' ,
furniture: 'Tables, chairs, sofas, beds, shelving, and other household furniture.' ,
clothing: 'Apparel, shoes, bags, and fashion accessories.' ,
sporting_goods: 'Bikes, fitness equipment, camping gear, and equipment for playing sports.' ,
toys_and_baby: 'Toys, games, strollers, car seats, cribs, and other children’s items.' ,
} as const ;
export type Category = keyof typeof CATEGORIES ;
export const CONDITIONS = [ 'for_parts' , 'fair' , 'good' , 'like_new' , 'new' ] as const ;
export type Condition = ( typeof CONDITIONS )[ number ];
export type Listing = {
id : string ;
title : string ;
description : string ;
category : Category ;
condition : Condition ;
priceUsd : number ;
photoCount : number ;
};
The category definitions are written as prose because Jev reads them from state when judging whether a listing belongs to the category. This is a general principle. If a judgment depends on a rule of yours, express the rule in the state in prose and reference it in your question, so the model uses your definition and not its own.
Hard rules are standard code and run before any call to Jev.
// rules.ts
import type { Category, Listing } from './listing' ;
// Typical sale price per category, in USD. In production this comes from your own sales data.
const TYPICAL_PRICE_USD : Record < Category , number > = {
electronics: 180 ,
furniture: 120 ,
clothing: 35 ,
sporting_goods: 90 ,
toys_and_baby: 40 ,
};
export type PriceSignal = 'far_below_typical' | 'typical' | 'far_above_typical' ;
export function priceSignal ( listing : Listing ) : PriceSignal {
const ratio = listing.priceUsd / TYPICAL_PRICE_USD [listing.category];
if (ratio < 0.2 ) return 'far_below_typical' ;
if (ratio > 10 ) return 'far_above_typical' ;
return 'typical' ;
}
// Hard rules. Anything here is a fact your code already knows, so no model is involved.
export function hardRuleFailures ( listing : Listing ) : string [] {
const failures : string [] = [];
if (listing.photoCount < 1 ) failures. push ( 'at least one photo is required' );
if (listing.priceUsd < 0.01 || listing.priceUsd > 50_000 ) failures. push ( 'price must be between $0.01 and $50,000' );
if (listing.title. trim (). length < 8 ) failures. push ( 'title must be at least 8 characters' );
if (listing.description. trim (). length < 30 ) failures. push ( 'description must be at least 30 characters' );
return failures;
}
Price bounds show the split inside a single check. The bounds are implemented in code, but a handbag listed for $120 in a category that usually sells for $1,800 points to counterfeit concerns, and interpreting that signal is a judgment. The code does the comparison and sends the finished label to Jev.
Structure the state Jev sees
State is the input half of the interface between your code and Jev. TypeSafe’s state docs describe it as what a panel of experts would receive before being asked for a judgment. Jev knows nothing about the subject except through the state, and every question in the request sees the same state.
Three rules for building the state apply in any domain. Use descriptive names for your fields, because the questions reference fields by path and the names convey meaning to the model. Include only the information the questions need, since irrelevant context distracts the model.
Pass computed facts as finished labels. Jev can use price_signal: 'far_below_typical' directly, while the raw values price_usd: 120 and typical_price: 1800 hand it a division problem.
Here is the state for one marketplace listing.
{
"listing" : {
"title" : "Louis Vuitton Neverfull MM tote" ,
"description" : "Mirror quality 1:1, same factory as the boutique version. Nobody can tell the difference. Text me on WhatsApp for more photos and a better price." ,
"category" : "clothing" ,
"declared_condition" : "new" ,
"price_usd" : 120
},
"category_definition" : "Apparel, shoes, bags, and fashion accessories." ,
"price_signal" : "typical"
}
Seller ID, timestamps, and photo URLs are omitted because none of the five questions use them. category_definition is included because the category-fit question compares the listing against it, and price_signal because the counterfeit criterion reads it.
Choose a question type for each judgment
Each Jev question has a type, and the type should follow from what the answer means. It fixes the shape of the answer, and the shape determines what you can do with it later.
Use choice when one option from a set you define should win and your code needs to know which one. You get the winner and a probability for every option, so you can see how much probability landed on the runners-up. Use noul when the answer is a proposition that is either true or false, and you get a single probability to threshold.
You can use a score to represent a degree. It is along an ordered scale. The result is a number which you can subtract, compare with other numbers, and which gives you the measure of the differences between them.
The criteria format varies by question type. choice questions receive an object that maps each option name to its description, with a maximum of 255 options. noul questions receive an optional object with true and false descriptions. An ideal question is concise and direct, so many questions don’t even need a description object. score questions receive an ordered array of level descriptions (with between 2 and 10 levels), and the answer is measured on the zero-based index of that array.
The marketplace uses all three. Prohibited item is a choice over five named categories and a none option, because the policy must know which rule was broken (a weapon is different from a counterfeit, as are the messages they trigger and the reviewers they go to). Category fit, title and description contradiction, and off-platform contact are noul questions, since each is a yes/no proposition and the only number the policy needs is the probability. Described condition is a score over five levels, because the policy asks how far apart the declared and described conditions are, and a distance can only be measured on an ordered scale.
Independent questions belong in one request. Jev evaluates them in parallel against the same state, so five questions cost one round trip, and a listing that trips two of them (the tote later in this post does) comes back with two distinct reasons. A broad question hides multiple judgments within one answer. Breaking broad questions into narrow, atomic ones makes them much easier to inspect, tune, and combine in code.
Write Jev criteria that survive a literal reading
Instructions say what to decide, and criteria define the answers. Jev applies both literally. That literalism is part of what makes Jev predictable, and it means the criteria are doing most of the work.
A few rules hold when writing criteria for any domain. Describe the options using the vocabulary present in the input text, because a criterion that names the situation (“a shared account with login details sent after payment”) matches text that never uses your policy’s word for it. For each noul , describe both sides so that near-misses fall on the correct side. Give every choice an explicit none for when nothing fits, so the model isn’t forced to pick something arbitrary.
For a score , describe each level as a concrete situation, because the levels make up the scale. And a question ID like prohibited is the key your code reads the answer back by, so don’t expect the model to interpret it as an instruction.
The client, condition scale, and prohibited items criteria are below. The noul criteria occupy the request itself. Look further for them in the next section.
// judge.ts
import { OpenRouter } from '@openrouter/sdk' ;
import { CATEGORIES, CONDITIONS, type Listing } from './listing' ;
import { priceSignal } from './rules' ;
const openrouter = new OpenRouter ({
apiKey: process.env. OPENROUTER_API_KEY , // server-side only
});
export const PROHIBITED = {
none: 'An ordinary secondhand or new item that a general marketplace allows.' ,
weapon_or_weapon_part:
'A firearm, firearm part or accessory, ammunition, stun gun, or a knife or tool marketed for fighting or self-defense.' ,
medication_or_medical_claim:
'Prescription medication, a controlled substance, or any product sold with a claim that it treats,
首次收录 · 2026-09-24 · 9.95 分