论文
arxiv:2609.39102
复制 markdown
虚假前沿:诊断与缓解自进化搜索智能体中的共作弊现象
发表于 9月30日
·
提交者
MEIJIA CHEN
于 10月1日
·
罗格斯大学
作者:
Meijia Chen
,
Hao Li
,
Zheng Lu
,
Hongshan Lin
,
Junbai Tian
,
Yichen Liu
,
Zijun Tian
,
Yufan Zou
,
Shuhan Sun
,
Hanxin Chen
,
Zeyu Zhang
,
Weizhi Du
,
Yueting Li
,
Tianyu Shi
,
Alaa Khamis
摘要
自进化搜索智能体通过联合优化生成问题的提议者(proposer)和回答问题的求解者(solver),构建自身的训练课程。这种闭环引入了一种我们称为“共作弊”(co-cheating)的故障模式:提议者和求解者在共享错误上日益趋同,导致内部奖励提升,但外部正确性并未相应提高。针对源证据的事后审计显示,随着自进化轮次的增加,共作弊现象愈发严重,即使循环内的训练信号改善,伪标签的正确性仍停滞或下降。最直接的缓解措施是在训练前验证提议:我们引入了多样本验证(MSV),即向同一模型在有源文档和无源文档的情况下各查询三次,以决定任务准入并替换不可靠的伪标签。MSV部分减少了虚假共识,但仍留下大量残留的共作弊现象,且每个候选项需额外消耗六次标注器生成次数。这些局限性催生了我们的主要方法 CrossFit:它将提议者的源文档划分为 A 组和 B 组;由 A 生成的问题由仅用 B 训练的辅助求解者评分,反之亦然。交叉拟合的共识决定了提议者的奖励,因此通过反馈求解器无法复现同源的伪标签,而原始求解者的更新规则保持不变。在 Qwen3.5-4B 和 Qwen3.5-9B 上重新运行循环,MSV 将虚假共识质量从 6.1% 降至 5.7%,从 8.8% 降至 7.2%,而 CrossFit 将其分别降至 3.0% 和 3.7%。使用排除源信息的反馈重放相同的提议,进一步将虚假共识降至 0.4% 和 0.1%,从而将反馈谱系与课程变化隔离开来。在七个下游搜索基准测试中,CrossFit 使平均性能比标准耦合自进化提高 8.8 和 8.4 分,比 Search-R1 提高 8.7 和 7.8 分(在 4B 和 9B 模型上)。
查看 arXiv 页面
查看 PDF
添加到收藏
社区
chenmeijia30
论文提交者
大约 16 小时前
TL;DR:在自进化搜索智能体中(提议者根据源文档编写问题和伪标签,求解者回答这些问题,两者的共识作为训练奖励),提议者和求解者可能收敛于共享错误:内部奖励持续上升,而实际正确性停滞。我们称之为共作弊,通过基于源的事后审计进行测量,并使用 CrossFit 修复它。
CrossFit 将提议者的源文档分为两折,并用仅用另一折训练的辅助求解者对每折的问题进行评分,因此同源伪标签永远无法作为奖励被回声反馈回来。主求解者仍然在所有数据上训练;只有反馈路径发生变化。
结果(Dr. Zero 循环,Qwen3.5-4B / 9B):
共作弊随自进化轮次增长:虚假共识质量达到 6.1% / 8.8%。
多样本验证(MSV)仅将其削减至 5.7% / 7.2%,每个候选项需额外消耗六次标注器生成次数。
CrossFit 将其削减至 3.0% / 3.7%;使用排除源信息的反馈重放相同提议,将其驱动至 0.4% / 0.1%。
下游应用:在七个搜索 QA 基准测试中,平均性能比耦合自进化提高 +8.8 / +8.4 分,比 Search-R1 提高 +8.7 / +7.8 分。
要点:当模型对自己进行评分时,应追踪评分者监督的来源,而不仅仅是标签看起来是否正确。
查看翻译
回复
编辑
预览
通过拖拽、粘贴或点击此处上传图片、音频和视频。
评论
· 注册或登录以发表评论
赞成
182
+170
在您的智能体中获取此论文:
hf papers 阅读 2609.39102
尚未安装最新的命令行界面工具?
引用该论文模型数量
0
无链接此论文的模型
在模型的 README.md 中引用 arxiv.org/abs/2609.39102,即可从本页面链接至该论文。
引用该论文数据集数量
0
无链接此论文的数据集
在数据集的 README.md 中引用 arxiv.org/abs/2609.39102,即可从本页面链接至该论文。
引用该论文 Spaces 数量
0
无链接此论文的 Space
在 Space 的 README.md 中引用 arxiv.org/abs/2609.39102,即可从本页面链接至该论文。
包含该论文的合集数量
3
Agent
合集
161 项
•
约 11 小时前更新
•
14
Code
合集
85 项
•
约 11 小时前更新
•
7
Search
合集
23 项
•
约 11 小时前更新
Papers
arxiv:2609.39102
Copy markdown
False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
Published on Sep 30
·
Submitted by
MEIJIA CHEN
on Oct 1
·
Rutgers University
Authors:
Meijia Chen
,
Hao Li
,
Zheng Lu
,
Hongshan Lin
,
Junbai Tian
,
Yichen Liu
,
Zijun Tian
,
Yufan Zou
,
Shuhan Sun
,
Hanxin Chen
,
Zeyu Zhang
,
Weizhi Du
,
Yueting Li
,
Tianyu Shi
,
Alaa Khamis
Abstract
Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer and solver increasingly agree on shared errors, so internal reward improves without a matching gain in external correctness. A post-hoc audit against source evidence shows co-cheating growing more severe over successive rounds of self-evolution, with pseudo-label correctness stagnating or declining even as the in-loop training signal improves. The most direct mitigation is to verify proposals before training: we introduce multi-sample verification (MSV), which queries the same model three times with the source and three times without it to decide task admission and replace unreliable pseudo-labels. MSV partially reduces false agreement but leaves substantial residual co-cheating and costs six extra labeler generations per candidate. These limitations motivate CrossFit, our main method: it partitions the proposer's source documents into groups A and B; questions generated from A are scored by an auxiliary solver trained only on B, and vice versa. The cross-fitted agreement determines proposer reward, so a same-source pseudo-label cannot be reproduced through the feedback solver, while the original solver's update rule is unchanged. Rerunning the loop with Qwen3.5-4B and Qwen3.5-9B, MSV reduces false-agreement mass from 6.1% to 5.7% and from 8.8% to 7.2%, whereas CrossFit reduces it to 3.0% and 3.7%. Replaying identical proposals with source-excluded feedback further reduces false agreement to 0.4% and 0.1%, isolating feedback ancestry from curriculum changes. Across seven downstream search benchmarks, CrossFit improves average performance over standard coupled self-evolution by 8.8 and 8.4 points and over Search-R1 by 8.7 and 7.8 points at 4B and 9B.
View arXiv page
View PDF
Add to collection
Community
chenmeijia30
Paper submitter
about 16 hours ago
TL;DR: In self-evolving search agents (a proposer writes questions + pseudo-labels from source documents, a solver answers them, and their agreement is the training reward), the proposer and solver can converge on shared errors: internal reward keeps rising while real correctness stalls. We call this co-cheating, measure it with a source-grounded post-hoc audit, and fix it with CrossFit.
CrossFit splits the proposer's source documents into two folds and scores each fold's questions with an auxiliary solver trained only on the other fold, so a same-source pseudo-label can never be echoed back as reward. The main solver still trains on everything; only the feedback path changes.
Results (Dr. Zero loop, Qwen3.5-4B / 9B):
Co-cheating grows over rounds of self-evolution: false-agreement mass reaches 6.1% / 8.8%.
Multi-sample verification (MSV) only trims it to 5.7% / 7.2%, at 6 extra labeler generations per candidate.
CrossFit cuts it to 3.0% / 3.7%; replaying identical proposals with source-excluded feedback drives it to 0.4% / 0.1%.
Downstream: +8.8 / +8.4 average points over coupled self-evolution and +8.7 / +7.8 over Search-R1 across seven search QA benchmarks.
Takeaway: when a model grades itself, track where the grader's supervision came from, not just whether the labels look right.
See translation
Reply
Edit
Preview
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Comment
· Sign up or log in to comment
Upvote
182
+170
Get this paper in your agent:
hf papers read 2609.39102
Don't have the latest CLI?
Models citing this paper
0
No model linking this paper
Cite arxiv.org/abs/2609.39102 in a model README.md to link it from this page.
Datasets citing this paper
0
No dataset linking this paper
Cite arxiv.org/abs/2609.39102 in a dataset README.md to link it from this page.
Spaces citing this paper
0
No Space linking this paper
Cite arxiv.org/abs/2609.39102 in a Space README.md to link it from this page.
Collections including this paper
3
Agent
Collection
161 items
•
Updated about 11 hours ago
•
14
Code
Collection
85 items
•
Updated about 11 hours ago
•
7
Search
Collection
23 items
•
Updated about 11 hours ago
首次收录 · 2026-10-02 · 10.95 分