微软 ThinkingBox 根据智能体留下的记录而非生成的语句对其进行评分,并进一步考察其能否连续二十次成功完成。目前该工具已通过 Hugging Face 提供。
图 1:ThinkingBox 让智能体在隔离的 MCP 工具会话中运行,随后对其终端后端状态及遗留的副作用进行评分。摘自我们的 ThinkingBox 论文。
本文由微软与 Hugging Face 联合发布,特别感谢 Tommy Guy(Enderis AI 创始人,前微软员工)、Hugging Face 的 Sergio Paniego,以及我们的前实习生 Zhuochun Li(匹兹堡大学)、Ali Keramati(加州大学尔湾分校)和 Youngmin Ko(西北大学)在共同撰写/审阅方面所做的贡献。
一位客户发来投诉。她那台价值 745 美元的厨房电器在纳什维尔配送中心被卡在快递“异常”状态,比预计送达日期晚了十五天。
AI 智能体进行了细致的工作。它调用了九次工具:拉取订单、检查物流追踪、查询她的客户档案、两次搜索退款政策、确认不存在工单后新建一个、记录时间线,并正确阅读了相关政策;她所在的账户层级确实不符合延迟送达补偿资格。
随后,它将工单标记为已解决,并回复:“既然您的问题已解决,还有什么我可以帮您的吗?”
有两处错误。首先,承运商的异常状态仍然开放,因此所需的最终状态处于“挂起”状态,等待解决。其次,客户从未得到对她实际问题的真实答复。
检查工具调用的 AI 评分器会看到九个格式良好的调用。检查智能体是否写入数据库的评分器也会看到这一点。真正产生分歧的是数据库本身。
这正是 ThinkingBox 所衡量的差距。在涵盖 507 个有状态业务流程的测试中,每个流程针对各种 LLM 模型运行 20 次,它根据终端后端状态和副作用对智能体进行评分。本文介绍了我们的发现、一致性所需的成本,以及如何通过 OpenEnv 自行运行该基准测试。
您可以亲自运行这一案例:上述示例改编自基准任务 sandbox_external_retail_group1.py:test_case_ST003_006,失败的执行检查仅涉及单个字段:工单状态显示为“已解决”,而所需的最终状态应为“挂起”。完整追踪记录见我们论文的附录 D.4 案例 3。
目录
工具调用并非结果
一次成功不等于可靠性
你能依赖智能体背后的模型吗?
一致性成本
失败特征
工作原理
自行运行
后续方向
想在阅读结果前尝试一下吗?请跳至“自行运行”部分。
工具调用并非结果
最终回复和有效的工具调用仅是代理指标。智能体可能听起来正确,却留下了错误的值、更改了错误的记录或产生了额外的副作用。只有它留下的记录才能定论。
这一差距十分显著。在涵盖 12 个 LLM 模型的 121,680 次有效试验的常见集合消融实验中,有 79,853 次尝试未能通过执行检查。在这些失败案例中,67.24% 的案例仍干净地终止、调用了状态变更工具并报告无最终工具错误。尽管如此,执行检查仍在其中 77.61% 的案例中发现字段值错误,在 43.30% 的案例中发现非预期的额外效果,在 25.36% 的案例中发现缺失的必要效果。这些状态检查结果存在重叠。
轨迹是一种主张。数据库状态是证据。重复性是信任测试。
一次成功不等于可靠性
一个智能体若正确执行了一次退款处理,却在接下来的四次中处理不当,它就不是一个可用的退款智能体。因此,每个任务均独立运行 20 次,每次从相同的干净后端开始,我们报告三项不同的指标:
表 1:我们报告的三个数字及其各自回答的问题。
指标
衡量内容
回答问题
pass@1
所有尝试中成功的比例
它通常表现如何?
pass@20
在 20 次尝试中至少成功一次的任务比例
它能否做到这一点?广度。
观察到的 20/20
在所有 20 次记录尝试中均通过的任务
它是否总能保持正确?
在本博客文章中,我们使用“观察到的 20/20”这一术语来字面统计在 507 项任务中有多少项实现了 20 次尝试全部通过。没有估算器,也没有平滑处理。
先从熟悉的视角开始。下表报告了 pass@1(单次尝试得分估计值),并按领域进行了细分。这是大多数排行榜发布的数据,单独来看,它就像是一个普通的能力排名。
表 2:ThinkingBox-Bench 按领域划分的 pass@1 (%)。每个模型在每项任务上进行了 20 次重复试验。粗体表示组内领先者;下划线表示亚军。单次尝试得分估计值的标准误差见我们 ThinkingBox 论文中的表 4。
模型
零售 (98)
车险 (100)
旅游 (104)
新银行 (104)
咨询 (101)
总体,任务加权 (507)
专有模型
Claude Opus 5.5
80.97
68.40
54.28
71.25
61.58
67.16
Claude Opus 5
80.71
65.80
49.95
70.62
66.19
66.50
GPT-5.4
76.33
62.65
68.12
65.34
54.60
65.36
GPT-5.6 Sol
67.65
65.30
60.34
59.09
57.52
61.91
Claude Sonnet 4.6
72.35
54.40
58.94
56.39
54.31
59.19
GPT-6 Astra
71.73
46.55
55.87
60.87
56.83
58.31
GPT-5.2
70.20
22.40
53.70
51.15
34.06
46.28
Claude Opus 4.6
68.62
8.30
21.11
35.67
27.82
32.09
o3-pro
37.70
2.95
17.31
24.28
14.60
19.31
Grok-4.3
43.93
2.60
15.14
1.78
9.55
14.38
开源模型
Kimi-K3
82.24
50.80
61.83
41.35
51.63
57.37
Qwen3.8-27B
64.03
47.85
53.41
47.88
45.69
51.70
DeepSeek-V4-Pro
68.21
29.65
43.13
44.86
31.04
43.26
Kimi-K2.6
53.72
24.50
39.52
33.65
37.33
37.66
GLM-5.1
58.67
25.70
35.43
13.27
34.06
33.19
Qwen3.6-27B
43.11
29.00
46.39
27.84
18.37
32.94
Qwen3.5-9B
19.90
0.70
4.71
1.15
2.33
5.65
Mistral-Large-3
11.28
1.30
8.99
1.15
0.74
4.66
Claude Opus 5.5 以 67.16% 的总成绩领先,比 Claude Opus 5 高出三分之二个百分点。Kimi-K3 是最强的开源模型,与 GPT-6-Astra 仅相差一个百分点。领域同样重要:Claude Opus 4.6 在零售领域的得分为 68.62%,但在车险领域仅为 8.30%。
一次成功的运行只能告诉你模型有能力完成工作。它并不能告诉你模型是否会再次成功。因此,让每项任务运行 20 次,并询问有多少得分得以保留。
图 2:每项模型的单次尝试得分在 20 次重复中保留了多少比例。
只有三个模型保留了大部分 pass@1 得分:GPT-6 Astra 保留了其单次尝试率的 78%,而 Claude Opus 5.5 和 Claude Opus 5 各保留了 71%。在另一端,GLM-5.1、Kimi-K2.6 和 DeepSeek-V4-Pro 每个仅保留了约 8%。
模型一次能做什么与它每次都能做什么之间的差距,才是完整的故事。
你能依赖代理背后的模型吗?
图 3:广度与一致性背道而驰。展示了十八个模型中的十二个;为了清晰起见,省略了 pass@1 低于 33% 的六个模型。
Kimi-K3 拥有我们测试过的任何模型中最广泛的覆盖范围。它在至少一次尝试中解决了基准测试的 93.89%:507 项任务中有 476 项。只有 31 项任务完全难倒了它,这是所有模型中最低的失败数量。在零售工作流方面,它以 82.24% 的 pass@1 领先,高于所有专有模型。
Kimi-K3 的一致性也较差。507 项任务中仅有 68 项(13.41%)在 20 次尝试中全部成功。
Claude Opus 5 则与此相反。它在至少一次尝试中解决的任务较少(79.09%;有 106 项任务完全难倒了它),但在每次尝试中都完成了基准测试的 47.53%。
较新的模型并不能解决这个问题。Claude Opus 5.5 在每次尝试的平均得分上高于 Claude Opus 5,分别为 67.16% 对 66.50%,并且解决了更多至少一次通过的任务。它在所有 20 次尝试中通过的任务数量完全相同:241 项。 headline 准确率提高的半个百分点并没有带来任何额外的可靠性。
Kimi-K3 比 Opus 5 多解决了 75 项至少一次通过的任务。
Opus 5 比 Kimi-K3 多稳定地解决了 173 项任务。
如果你正在为涉及真实记录的工作选择模型,pass@20 不是你应该查看的那一列数据。
一致性成本几何
能力比较通常止步于分数。对于任何部署者而言,相关的问题是完成一个有效工作单元的成本是多少。我们将其衡量为每次成功任务尝试的成本。之所以称为“任务尝试”,是因为每个基准测试任务都会重复运行,且成本是按每次尝试产生的,因此 pass@1 是匹配的合格率分母。
我们从模型在完整的 507 × 20 次活动中的记录令牌使用量中获取数据,并按 OpenRouter + 上可用的无折扣目录价格进行定价,同时逆转促销折扣并排除声明了量化处理的端点。每个模型的输入、输出和缓存费率均来自单一提供商端点。
然后,我们将一次运行的成本除以成功的尝试次数:
每次成功任务尝试的成本 = 507 次尝试的预估成本(每项任务一次)÷ (507 × pass@1)
这是一个比较效率指数,而非发票,也不是服务单个生产请求的价格。它衡量的是单次成功的成本,而非一致性。接下来我们计算一致性的成本。
示例:GPT-5.4 在 507 次尝试(每项任务一次)中的成本为 43.49 美元,pass@1 为 65.36%,因此 43.49 美元 ÷ (507 × 0.6536) = 每次成功任务尝试 0.131 美元。
帕累托成本前沿
如果没有任何其他模型既不比它更贵且准确率至少相当,则该模型位于前沿上。有三个模型符合这一条件;其他所有模型至少在其中一个维度上处于劣势。
图 4:每次成功任务尝试的成本与 pass@1 的对比。带环的点为帕累托成本前沿模型。
前沿分为三个阶段。GPT-5.6 Sol 的成功成本最低,为 0.127 美元;GPT-5.4 使 pass@1 提高了 3.45 个百分点,每次成功的成本增加 0.004 美元;Claude Opus 5.5 又增加了 1.80 个百分点,每次成功的成本为 0.276 美元。这三个模型都保持在成本前沿线上,因为没有更便宜的模型能达到其 pass@1 水平。
Claude Opus 5 是最明显的案例:其每次成功尝试的成本为 0.475 美元,pass@1 为 66.50%,它既比 Claude Opus 5.5(成本 0.276 美元,pass@1 为 67.16%)更贵,准确率也更低。
现在计算一致性的成本
成功成本奖励的是便宜且经常正确的模型。它并不奖励每次都能正确的模型。因此,我们还计算每个可靠任务的成本:即整个 20 次运行活动的成本除以模型在全部 20 次尝试中都通过的任务数量。
每个可靠任务的成本 = 507 次尝试的 20 次运行的预估成本 ÷ 20/20 通过的任务数
示例:GPT-6-Astra 在该活动中的成本为 20 × 86.03 美元 = 1,720.60 美元,并在每次尝试中通过了 231 项任务,因此 1,720.60 美元 ÷ 231 = 每个可靠任务 7.45 美元。
表 3:至少有一个观察到 20/20 通过任务的模型中,每个可靠任务成本最低的九个模型,按从低到高排序。预估美元金额,非实际云账单。
模型
20/20 通过的任务数
20 次运行的预估成本
每个可靠任务的成本
GPT-5.4
128 (25.25%)
869.80 美元
6.80 美元
GPT-6 Astra
231 (45.56%)
1,720.60 美元
7.45 美元
Claude Opus 5.5
241 (47.53%)
1,880.77 美元
7.80 美元
GPT-5.6 Sol
82 (16.17%)
800.00 美元
9.76 美元
Claude Opus 5
241 (47.53%)
3,206.00 美元
13.30 美元
Claude Sonnet 4.6
102 (20.12%)
1,587.60 美元
15.56 美元
GPT-5.2
44 (8.68%)
878.00 美元
19.95 美元
Kimi-K3
68 (13.41%)
1,406.40 美元
20.68 美元
Qwen3.8-27B
38 (7.50%)
925.80 美元
24.36 美元
现在按一致性进行排名。GPT-5.4 最便宜,为 6.80 美元,尽管只有 128 项任务达到标准。GPT-6 Astra 在 7.45 美元的成本下达到 231 项,Claude Opus 5.5 并列最高,达到 241 项,成本为 7.80 美元。
这三者之间没有一方完全主导另一方:每增加一个可靠任务都需要更高的成本。Claude Opus 5 也通过了 241 项任务,但成本为 13.30 美元,因此 Opus 5.5 完全优于它。GPT-5.6 Sol 单次成功的成本最低,为 0.127 美元,但每个可靠任务的成本为 9.76 美元。获得正确答案的最便宜方式并不是获得可靠答案的最便宜方式。
失败特征
我们为每次失败的轨迹分配一个确定性的诊断特征,且主要结论具有可操作性:大约五分之四的失败源于工具处理,而非推理能力。根据我们论文中表 5 的消融研究:
失败特征
失败占比
工具使用
79.9%
状态更新错误
10.3%
用户解决方案不完整
7.0%
无状态变更操作
2.9%
这些是各模型份额和可观察标签的未加权平均值,而非唯一的因果解释。
实际模式很简单:代理通常能推进到足以尝试工作流的程度,然后因工具错误、前置条件失败或空查询而无法恢复。这首先是一个重试和错误恢复问题,其次才是模型问题。
难度也因领域而异:在上述表 2 列出的模型中,零售业的 pass@1 平均通过率为 59.52%,而汽车保险的平均通过率为 33.83%。
该怎么办。将 20/20 通过率视为设计输入,而非最终判决。基准测试所评估的同一信号在生产环境中同样可用:在提交之前检查终端状态,而不是依赖模型的摘要。
对工具和系统错误进行分类,以便重试针对可恢复的错误。缩减工具表面范围至工作流所需的部分。对于无法低成本撤销的变更,要求人工批准。我们尚未测量这些措施在该基准测试中的提升效果,而这正是该环境现在使其可测试的一类事情。
工作原理
ThinkingBox 是代理沙箱,而 ThinkingBox-Bench 是用于评估代理的数据集基准。本文顶部的图表展示了循环过程;以下是各部分的功能说明。
图 5:上述图 1 面板 A 中的沙箱循环:隔离的工具会话、终端数据库状态、副作用、可执行裁判。
每个任务定义了一个起始后端状态、用户目标、可用的 MCP 工具、领域策略以及对终端状态的 executable checks(可执行检查)。模拟用户持有私有上下文(预订参考号、偏好或出生日期),仅在被要求时才释放这些信息。
每次尝试都会获得一个隔离的 MCP 会话,并初始化全新的状态。同一任务的两次尝试永远不会共享数据库行或缓存的工具状态,这使得 20 次试验的比较具有意义。
最后,副作用提取器会推导实际发生了哪些变更,确定性裁判将其与所需的最终状态进行比较,接受产生正确结果的任何轨迹,同时拒绝错误、缺失或多余的副作用。对于没有清晰数据库值的要求(“代理是否披露了这不是保证的?”),通过狭窄的二元评分标准问题来处理语义。507 个任务中有 477 个仅基于状态进行评分;另外 30 个增加了响应评分标准。
信任边界:模型可以看到任务、对话和工具模式。黄金状态、断言、评分内部逻辑和凭据保留在评估器端。
自行运行
ThinkingBox 现已在 Hugging Face 上发布,包括框架和数据集。ThinkingBox-Bench 现在位于 OpenEnv 接口之后,每个完成的剧集都会返回一个二元的通过/失败奖励。发布的适配器专为评估而设计;在非基准测试场景中,训练工作流可以使用相同的接口。
开始之前
已在 Linux 和 WSL 上测试,使用 Python 3.11+、uv 和 Docker。你还需要在固定版本处签出 thinkingbox-data,并为代理、模拟用户和裁判提供模型端点。一个端点可以服务所有这三个角色,这是最简单的启动方式。OpenEnv 镜像仅启动 OpenEnv API;其他所有内容均由你自己运行。
安装
git clone https://github.com/huggingface/OpenEnv
cd OpenEnv
uv sync --project envs/thinkingbox_env --frozen
git clone https://github.com/microsoft/thinkingbox-data
git -C thinkingbox-data checkout thinkingbox-bench-v1.0
uv tool install "thinkingbox @ git+https://github.com/microsoft/thinkingbox"
启动 Typesense
在第二个终端中,启动 Typesense 30.1 并等待其健康检查:
mkdir -p .typesense-data
docker run -- rm -d --name thinkingbox-typesense \
-p 8108:8108 \
-v " $PWD /.typesense-data:/data" \
typesense/typesense:30.1 \
--data-dir /data --api-key=Fake --enable-cors
until curl -fsS http://127.0.0.1:8108/health; do sleep 1; done
启动 MCP 服务器
在第三个终端中,启动 Session Proxy 和 MCP 服务器。
cd OpenEnv
tb mcp-start --host 127.0.0.1 --port 7111 \
--servers "$PWD/thinkingbox-data/servers/servers.yaml"
curl -fsS http://127.0.0.1:7111/health
启动 OpenEnv 服务器
回到第一个终端,启动 O
Microsoft ThinkingBox grades AI agents on the records they leave behind, not the sentences they generate, and then asks whether they can do it twenty times in a row. It is now available through Hugging Face.
Figure 1: ThinkingBox runs an agent against isolated MCP tool sessions, then grades the terminal backend state and side effects it leaves behind. From our ThinkingBox paper .
This is a joint blog by Microsoft and Hugging Face, special thanks to Tommy Guy (founder at Enderis AI, previously Microsoft), Sergio Paniego from Hugging Face and our former interns Zhuochun Li (University of Pittsburgh), Ali Keramati (UC Irvine), Youngmin Ko (Northwestern) for co-authoring/reviewing efforts.
A customer writes in. Her $745 kitchen appliance has been stuck in a courier "exception" at a Nashville distribution center, fifteen days past its estimated delivery date.
The AI agent does careful work. Nine tool calls: it pulls the order, checks tracking, looks up her customer profile, searches the refund policy twice, confirms no ticket exists, opens one, documents the timeline, and reads the policy correctly; her account segment genuinely does not qualify for late-delivery compensation.
Then it closes the ticket as resolved and replies “ Since your query is resolved, is there anything I may assist you with? ”
Two things are wrong. The carrier exception is still open, so the required end state was on hold , pending resolution. And the customer never got a real answer to what she actually asked.
An AI grader checking tool calls would see nine well-formed ones. The grader checking whether the agent wrote to the database would see that too. The database is what disagrees .
That gap is what ThinkingBox measures. Across 507 stateful business workflows, each run 20 times against various LLM models, it grades agents on terminal backend state and side effects. This post covers what we found, what consistency costs, and how to run the benchmark yourself through OpenEnv .
You can run this one yourself: the example above is adapted from a benchmark task sandbox_external_retail_group1.py:test_case_ST003_006 , and the executable check that fails is a single field: the ticket's status is solved where the required end state is hold. The full trace is in Appendix D.4, Case 3 of our paper .
Contents
A tool call is not an outcome
One success is not reliability
Can you depend on the model behind your agent?
What consistency costs
Failure signatures
How it works
Run it yourself
Where this goes next
Want to try it before reading the results? Skip to section Run it yourself .
A tool call is not an outcome
Final responses and valid tool calls are only proxies. An agent can sound correct while leaving the wrong value, changing the wrong record, or creating an extra side effect. Only the records it leaves behind settle the question.
The gap is substantial. In a common-set ablation covering 121,680 valid trials across 12 LLM models, 79,853 attempts failed the executable checks. Of those failures, 67.24% still terminated cleanly, invoked a state-changing tool, and reported no final tool error. Executable checks nevertheless found wrong field values in 77.61% of them, unintended extra effects in 43.30%, and missing required effects in 25.36%. Those state-check findings overlap.
A trajectory is a claim. Database state is the evidence. Repetition is the trust test.
One success is not reliability
An agent that processes a refund correctly once and mishandles it the next four times is not a working refund agent. So every task runs 20 independent times , each from an identical clean backend, and we report three different things:
Table 1: The three numbers we report, and the question each one answers.
Metric
What it measures
What it answers
pass@1
Share of all attempts that succeeded
How does it usually do?
pass@20
Share of tasks solved at least once in 20 tries
Can it ever do this? Breadth.
Observed 20/20
Tasks that actually passed all 20 recorded attempts
Can it always be correct?
We use observed 20/20 in this blog post as the literal count of how many of the 507 tasks passed 20 out of 20. No estimator, no smoothing.
Starting with the familiar view. The table below reports pass@1, the single-attempt score estimate, broken out by domain. This is the number most leaderboards publish, and on its own it reads like an ordinary capability ranking.
Table 2: ThinkingBox-Bench pass@1 (%) by domain. Each model is evaluated on every task for 20 repeated trials. Bold marks the group leader; underline marks the runner-up. The standard errors for the single attempt score estimates are provided in Table 4 in our ThinkingBox paper .
Model
Retail (98)
Auto insurance (100)
Travel (104)
Neobank (104)
Consulting (101)
Overall, task-weighted (507)
Proprietary models
Claude Opus 5.5
80.97
68.40
54.28
71.25
61.58
67.16
Claude Opus 5
80.71
65.80
49.95
70.62
66.19
66.50
GPT-5.4
76.33
62.65
68.12
65.34
54.60
65.36
GPT-5.6 Sol
67.65
65.30
60.34
59.09
57.52
61.91
Claude Sonnet 4.6
72.35
54.40
58.94
56.39
54.31
59.19
GPT-6 Astra
71.73
46.55
55.87
60.87
56.83
58.31
GPT-5.2
70.20
22.40
53.70
51.15
34.06
46.28
Claude Opus 4.6
68.62
8.30
21.11
35.67
27.82
32.09
o3-pro
37.70
2.95
17.31
24.28
14.60
19.31
Grok-4.3
43.93
2.60
15.14
1.78
9.55
14.38
Open-weight models
Kimi-K3
82.24
50.80
61.83
41.35
51.63
57.37
Qwen3.8-27B
64.03
47.85
53.41
47.88
45.69
51.70
DeepSeek-V4-Pro
68.21
29.65
43.13
44.86
31.04
43.26
Kimi-K2.6
53.72
24.50
39.52
33.65
37.33
37.66
GLM-5.1
58.67
25.70
35.43
13.27
34.06
33.19
Qwen3.6-27B
43.11
29.00
46.39
27.84
18.37
32.94
Qwen3.5-9B
19.90
0.70
4.71
1.15
2.33
5.65
Mistral-Large-3
11.28
1.30
8.99
1.15
0.74
4.66
Claude Opus 5.5 leads overall at 67.16%, two-thirds of a point above Claude Opus 5. Kimi-K3 is the strongest open-weights model , within a point of GPT-6-Astra. Domain matters just as much : Claude Opus 4.6 scores 68.62% on retail but 8.30% on auto insurance.
One good run tells you a model can do the work. It does not tell you whether it will do it again. So run every task 20 times and ask how much of that score survives.
Figure 2: How much of each model's single-attempt score survives 20 repeats.
Only three hold on to most of their pass@1 scores: GPT-6 Astra retains 78% of its single-attempt rate, and Claude Opus 5.5 and Claude Opus 5 each retain 71%. At the other end, GLM-5.1, Kimi-K2.6 and DeepSeek-V4-Pro each keep about 8%.
The gap between what a model can do once and what it does every time is the whole story.
Can you depend on the model behind your agent?
Figure 3: Breadth and consistency pull apart. Twelve of the eighteen models are shown; six below 33% pass@1 are omitted for legibility.
Kimi-K3 has the broadest coverage of any model we tested. It solves 93.89% of the benchmark at least once: 476 of 507 tasks. Only 31 tasks defeat it entirely, the lowest count in the field. On retail workflows it leads outright at 82.24% pass@1, ahead of every proprietary model.
Kimi-K3 is also among the least consistent. Just 68 of 507 tasks, 13.41%, succeed in all 20 attempts.
Claude Opus 5 inverts this. It solves fewer tasks at least once (79.09%; 106 defeat it entirely) but completes 47.53% of the benchmark on every single attempt.
A newer model does not fix this. Claude Opus 5.5 scores higher than Claude Opus 5 on every-attempt average, 67.16% against 66.50%, and solves more tasks at least once. It passes exactly the same number of tasks on all 20 attempts: 241. Half a point of headline accuracy bought no additional dependability at all.
Kimi-K3 solves 75 more tasks at least once than Opus 5.
Opus 5 solves 173 more tasks consistently than Kimi-K3.
If you are choosing a model for work that touches real records, pass@20 is the wrong column to look at.
What consistency costs
Capability comparisons usually stop at the score. For anyone deploying, the relevant question is what a successful unit of work costs. We measure that as cost per successful task attempt. We say task attempt because every benchmark task is run repeatedly and cost is incurred per attempt, so pass@1 is the matching quality denominator.
We took each model's recorded token usage from its full 507 × 20 campaign and priced it at undiscounted list rates available on OpenRouter + , reversing promotional discounts and excluding endpoints that declare quantization. Input, output and cache rates all come from one provider endpoint per model.
Then we divided one run's cost by the number of attempts that succeeded :
Cost per successful task attempt = estimated cost for 507 attempts, one per task ÷ (507 × pass@1)
This is a comparative efficiency index, not an invoice, and not the price of serving one production request. It also prices single successes, not consistency. We price consistency next.
Example: GPT-5.4 costs $43.49 for 507 attempts (one attempt per task) and has 65.36% pass@1, so $43.49 ÷ (507 × 0.6536) = $0.131 per successful task attempt.
Pareto cost frontier
A model is on the frontier if no other model is both no more expensive and at least as accurate . Three models qualify; every other model is dominated on at least one axis.
Figure 4: Cost per successful task attempt against pass@1. Ringed dots are pareto cost frontier models.
The frontier has three steps. GPT-5.6 Sol has the lowest cost per success at $0.127; GPT-5.4 raises pass@1 by 3.45 percentage points for $0.004 more per success; Claude Opus 5.5 adds another 1.80 points at $0.276 per success. Each of the three remains on the cost frontier line because no cheaper model matches its pass@1.
Claude Opus 5 is the clearest case: at $0.475 per successful attempt and 66.50% pass@1, it is both more expensive and less accurate than Claude Opus 5.5 at $0.276 and 67.16%.
Now price consistency
Cost per success rewards a model that is cheap and often right. It does not reward a model that is right every time. So we also compute cost per dependable task: the cost of the full 20-run campaign divided by the number of tasks the model passed on all 20 attempts.
Cost per dependable task = estimated cost of 20 runs of 507 attempts ÷ tasks passing 20/20
Example: GPT-6-Astra costs 20 × $86.03 = $1,720.60 for the campaign and passes 231 tasks on every attempt, so $1,720.60 ÷ 231 = $7.45 per dependable task.
Table 3: The nine lowest costs per dependable task among models with at least one observed 20/20 task, sorted low to high. Estimated $, not actual cloud bills.
Model
Tasks passing 20/20
Est. cost, 20 runs
Cost per dependable task
GPT-5.4
128 (25.25%)
$869.80
$6.80
GPT-6 Astra
231 (45.56%)
$1,720.60
$7.45
Claude Opus 5.5
241 (47.53%)
$1,880.77
$7.80
GPT-5.6 Sol
82 (16.17%)
$800.00
$9.76
Claude Opus 5
241 (47.53%)
$3,206.00
$13.30
Claude Sonnet 4.6
102 (20.12%)
$1,587.60
$15.56
GPT-5.2
44 (8.68%)
$878.00
$19.95
Kimi-K3
68 (13.41%)
$1,406.40
$20.68
Qwen3.8-27B
38 (7.50%)
$925.80
$24.36
Now rank by consistency. GPT-5.4 is the cheapest at $6.80, though only 128 tasks meet the bar. GPT-6 Astra reaches 231 at $7.45, and Claude Opus 5.5 the joint-highest 241 at $7.80.
None of the three dominates the others: each additional dependable task costs more. Claude Opus 5 also passes 241, but at $13.30, so Opus 5.5 dominates it outright. GPT-5.6 Sol, the cheapest per single success at $0.127, costs $9.76 per dependable task. The cheapest way to get a right answer is not the cheapest way to get a dependable one.
Failure signatures
We assign each failed trace one deterministic diagnostic signature, and the headline is actionable: roughly four in five failures are tool handling, not reasoning. Across an ablation study in Table 5 of our paper :
Failure signature
Share of failures
Tool usage
79.9%
Wrong state updates
10.3%
Incomplete user resolutions
7.0%
No state-changing action
2.9%
These are unweighted averages of per-model shares and observable labels, not unique causal explanations.
The practical pattern is simple: agents usually get far enough to attempt the workflow, then fail to recover from tool errors, failed preconditions, or empty lookups. That is a retry and error-recovery problem before it is a model problem.
Difficulty also changes by domain: across the models listed in Table 2 above, retail averages 59.52% pass@1 while auto insurance averages 33.83%.
What to do about it. Treat the 20/20 rate as a design input, not a verdict. The same signal the benchmark grades on is available in production: check the terminal state before you commit, not the model's summary of it.
Classify tool and system errors so retries target the recoverable ones. Cut the tool surface to what the workflow needs. And require human approval on the changes you cannot cheaply reverse. We have not measured the lift from any of these on this benchmark, which is exactly the kind of thing the environment now makes testable.
How it works
ThinkingBox is the agent sandbox, while ThinkingBox-Bench is a dataset benchmark to evaluate agents. The diagram at the top of this post shows the loop; here is what each part does.
Figure 5: The sandbox loop from panel A in Figure 1 above: isolated tool session, terminal database state, side effects, executable judges.
Each task defines a starting backend state, a user goal, the available MCP tools, the domain policy, and executable checks over the terminal state. A simulated user holds private context (a booking reference, a preference, or a date of birth) and releases it only when asked.
Every attempt gets an isolated MCP session with freshly initialized state. Two attempts of the same task never share a database row or cached tool state, which is what makes 20-trial comparison meaningful.
At the end, a side-effect extractor derives what actually changed, and deterministic judges compare it against the required end state, accepting any trajectory that produces the right outcome while rejecting wrong, missing or extra effects. For requirements with no clean database value ("did the agent disclose this is not guaranteed?"), a narrow binary rubric question handles the semantics. 477 of the 507 tasks are graded on state alone; 30 add response rubrics.
The trust boundary: the model sees tasks, dialogue and tool schemas. Golden state, assertions, grading internals and credentials stay on the evaluator side.
Run it yourself
ThinkingBox is now on Hugging Face, both the harness and the dataset . ThinkingBox-Bench now sits behind the OpenEnv interface, and each finished episode returns a binary pass/fail reward. The released adapter is designed for evaluation; separate, non-benchmark scenarios can use the same interface in training workflows.
Before you start
Tested on Linux and WSL, with Python 3.11+, uv and Docker. You also need a thinkingbox-data checkout at the pinned release and model endpoints for the agent, simulated user and judge. One endpoint can serve all three roles, which is the simplest way to start. The OpenEnv image starts only the OpenEnv API ; everything else you run yourself.
Install
git clone https://github.com/huggingface/OpenEnv
cd OpenEnv
uv sync --project envs/thinkingbox_env --frozen
git clone https://github.com/microsoft/thinkingbox-data
git -C thinkingbox-data checkout thinkingbox-bench-v1.0
uv tool install "thinkingbox @ git+https://github.com/microsoft/thinkingbox"
Start Typesense
In a second terminal , start Typesense 30.1 and wait for its health check:
mkdir -p .typesense-data
docker run -- rm -d --name thinkingbox-typesense \
-p 8108:8108 \
-v " $PWD /.typesense-data:/data" \
typesense/typesense:30.1 \
--data-dir /data --api-key=Fake --enable-cors
until curl -fsS http://127.0.0.1:8108/health; do sleep 1; done
Start the MCP servers
In a third terminal , start the Session Proxy and MCP servers.
cd OpenEnv
tb mcp-start --host 127.0.0.1 --port 7111 \
--servers " $PWD /thinkingbox-data/servers/servers.yaml"
curl -fsS http://127.0.0.1:7111/health
Start the OpenEnv server
Back in the first terminal, start the O
首次收录 · 2026-10-05 · 9.91 分