论文
arxiv:2609.25804
复制 Markdown
《有品位的智能体:衡量与提升长程任务中的品味》
发布于 9月22日
·
提交者
Wenbo Pan
于 9月23日
今日第1篇论文
·
微软
作者:
Wenbo Pan
,
Zhichao Liu
,
Shujie Liu
,
Jingying Zeng
,
Chin-Yew Lin
,
Xianfeng Tang
,
Yan Lu
,
Qi He
,
Xiaohua Jia
摘要
大型语言模型(LLM)智能体越来越多地参与长程任务,而它们在过程中做出的决策——例如测试哪个假设或基于哪个实现进行构建——决定了整个运行过程的结果。做出良好的决策正成为工程和科研智能体的关键能力。我们将这种做出良好长程决策的能力称为智能体的“品味”(taste)。尽管现有的基准测试衡量的是智能体在长程任务上的端到端成功率,但没有任何一个基准测试衡量智能体的品味。为了解决这一问题,我们构建了 Taste-Bench,这是一个由工程和研究任务中智能体生成的轨迹自动构建的品味问题基准。每个问题呈现一个决策分叉点(decision fork),即轨迹中多个方向可供选择且其中一个能带来更好结果的节点,被评估的模型在不看到分叉后结果的情况下从这些方向中进行选择。我们无需人工标注,即可从同一任务的并行尝试以及单个轨迹内部的迂回路径中自动挖掘这些分叉点。我们在 Taste-Bench 上评估了前沿模型,发现表现最好的模型仅正确回答了 59.7% 的问题。我们进一步发现,对于所有模型而言,决策依据出现在轨迹较后位置的案例要困难得多,且增加推理预算并不能提高准确率。最后,我们证明品味是可以训练的。我们将一位知晓最终结果的“教师”模型的判断蒸馏到“学生”模型中,该学生在未见过的任务上做出了更好的决策,并在保留的 SWE-bench Pro 任务上提高了端到端的成功率。
查看 arXiv 页面
查看 PDF
GitHub
1
添加到合集
社区
wenbopan
论文作者
论文提交者
约 18 小时前
大家好,我是作者 👋 我们研究了 LLM 智能体的品味:即在结果不可见之前,智能体在决策分叉点选择更好方向的能力。
🔹 Taste-Bench:从 SWE-bench Pro 和 METR AI R&D 轨迹中自动挖掘的 502 个决策分叉点,其标签由后续实际发生的情况确定(与人工审查的一致性为 98.8%)。
🔹 14 个前沿模型中表现最好的得分 59.7%(随机猜测 = 25%)。随着决策依据向未来推移,准确率从 62.3% 降至 21.0%,且更大的推理预算无助于提升效果。
🔹 品味是可训练的:将拥有后见之明的教师模型进行蒸馏,使 Qwen3.6-27B 在未见过任务上的得分从 30.0% 提升至 47.9%,其建议也将 SWE-bench Pro 的成功率从 14.6% 提升至 33.7%。
数据集:https://huggingface.co/datasets/wenbopan/taste-bench · 代码:https://github.com/wbopan/tastebench
欢迎提问!
查看翻译
回复
O96a
约 10 小时前
当单次运行包含十几个决策点且奖励稀疏时,你究竟如何区分品味与运气?一条通过糟糕策略的幸运轨迹,在运行十次之前看起来与有品味的轨迹完全相同。我希望看到的是,品味得分是否比单纯计算失败次数更能预测最终结果——如果是这样,这就不再仅仅是一个基准测试,而变成了一个调试工具。我花了太多时间盯着失败的智能体运行,纠结于是策略错了还是运气不好。这正是这个研究可以填补的空白。
查看翻译
回复
wenbopan
论文作者
论文提交者
约 10 小时前
感谢您的提问!生成模型将被指示提供证据,证明目标决策确实有助于最终成功,否则该幸运轨迹将被丢弃。
查看翻译
回复
researchstudio-bot
约 9 小时前
这是 ResearchStudio 团队发送的自动消息。
我们为这篇论文制作了一个互动式 ResearchStudio Reel。它包含一张视觉海报、一段视频和一个博客,所有文件均可下载且支持编辑格式。
打开 ResearchStudio Reel →
从 Hugging Face 下载所有文件
如果您觉得这个 Reel 有帮助,请给这条评论点赞!
想探索或为更多论文创建 Reel?请访问 ResearchStudio 演示页面。
查看翻译
回复
akilez64
大约 4 小时前
•
此评论已被隐藏(标记为垃圾信息)
编辑
预览
通过拖拽、粘贴或点击此处上传图片、音频和视频。
评论
· 注册或登录以发表评论
赞同
111
+99
在您的智能体中获取这篇论文:
hf papers read 2609.25804
没有最新的 CLI?
引用此论文的模型
0
没有链接此论文的模型
在模型 README.md 中引用 arxiv.org/abs/2609.25804 以从本页面链接它。
引用此论文的数据集
1
wenbopan/taste-bench
查看器
•
大约 19 小时前更新
•
502
•
17
引用此论文的 Spaces
0
没有链接此论文的 Space
在 Space README.md 中引用 arxiv.org/abs/2609.25804 以从本页面链接它。
包含此论文的合集
0
没有包含此论文的合集
将此论文添加到合集中以从本页面链接它。
Papers
arxiv:2609.25804
Copy markdown
The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks
Published on Sep 22
·
Submitted by
Wenbo Pan
on Sep 23
·
Microsoft
Authors:
Wenbo Pan
,
Zhichao Liu
,
Shujie Liu
,
Jingying Zeng
,
Chin-Yew Lin
,
Xianfeng Tang
,
Yan Lu
,
Qi He
,
Xiaohua Jia
Abstract
LLM agents increasingly work on long-horizon tasks, and the decisions they make along the way, such as which hypothesis to test or which implementation to build on, determine the outcome of the whole run. Making these decisions well is becoming a key capability for both engineering and research agents. We refer to the ability to make good long-horizon decisions as the taste of an agent. While existing benchmarks measure the end-to-end success of agents on long-horizon tasks, none of them measures the taste of an agent. To address this problem, we build Taste-Bench, a benchmark of taste questions constructed automatically from trajectories that agents produced in engineering and research tasks. Each question presents a decision fork, a point in a trajectory where multiple directions are available and one of them leads to a better outcome, and the evaluated model chooses among these directions without seeing what happens after the fork. We mine these forks automatically from parallel attempts at the same task and from detours inside a single trajectory, without needing human annotation. We evaluate frontier models on Taste-Bench and find that the best model answers only 59.7% of the questions correctly. We further find that forks whose deciding evidence appears later in the trajectory are much harder for every model, and that a larger reasoning budget does not improve the accuracy. Finally, we show that taste can be trained. We distill the judgment of a teacher that has seen the outcome into a student model, and the student makes better decisions on unseen tasks and improves end-to-end success on held-out SWE-bench Pro tasks.
View arXiv page
View PDF
GitHub
1
Add to collection
Community
wenbopan
Paper author
Paper submitter
about 18 hours ago
Hi all, author here 👋 We study the taste of LLM agents: their ability to pick the better direction at a decision fork before the outcome is visible.
🔹 Taste-Bench: 502 decision forks mined automatically from SWE-bench Pro and METR AI R&D trajectories, labeled by what actually happened later (98.8% agreement with human review).
🔹 The best of 14 frontier models scores 59.7% (random = 25%). Accuracy drops from 62.3% to 21.0% as the deciding evidence moves further into the future, and a larger reasoning budget does not help.
🔹 Taste is trainable: distilling a hindsight teacher lifts Qwen3.6-27B from 30.0% to 47.9% on unseen tasks, and its advice raises SWE-bench Pro success from 14.6% to 33.7%.
Dataset: https://huggingface.co/datasets/wenbopan/taste-bench · Code: https://github.com/wbopan/tastebench
Happy to answer questions!
See translation
Reply
O96a
about 10 hours ago
How do you actually separate taste from luck when a single run has a dozen decision points and sparse reward? One lucky trajectory through a bad policy looks identical to a tasteful one until you've run it ten times. What I'd want to see is whether the taste score predicts final outcome better than just counting failures — if it does, this stops being a benchmark and becomes a debugging tool. I've spent too many hours staring at a failed agent run wondering if the policy was wrong or the dice just rolled badly. That's the gap this could actually fill.
See translation
Reply
wenbopan
Paper author
Paper submitter
about 10 hours ago
Thank you for your question! The generator model will be instructed to provide evidence of that the target decision indeed contributes to the final success, otherwise the lucky trajectory will be dropped.
See translation
Reply
researchstudio-bot
about 9 hours ago
This is an automated message from the ResearchStudio team.
We created an interactive ResearchStudio Reel for this paper. It includes a visual poster, a video, and a blog, all available for download in editable formats.
Open the ResearchStudio Reel →
Download all files from Hugging Face
Please give this comment a thumbs up if you find the Reel helpful!
Want to explore or create Reels for more papers? Visit the ResearchStudio demo.
See translation
Reply
akilez64
about 4 hours ago
•
This comment has been hidden (marked as Spam)
Edit
Preview
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Comment
· Sign up or log in to comment
Upvote
111
+99
Get this paper in your agent:
hf papers read 2609.25804
Don't have the latest CLI?
Models citing this paper
0
No model linking this paper
Cite arxiv.org/abs/2609.25804 in a model README.md to link it from this page.
Datasets citing this paper
1
wenbopan/taste-bench
Viewer
•
Updated about 19 hours ago
•
502
•
17
Spaces citing this paper
0
No Space linking this paper
Cite arxiv.org/abs/2609.25804 in a Space README.md to link it from this page.
Collections including this paper
0
No Collection including this paper
Add this paper to a collection to link it from this page.
首次收录 · 2026-09-24 · 10.95 分