论文
arxiv:2609.24972
复制 markdown
RRSI:智能体工具链的正则化递归自改进
发布于 9月21日
·
提交者
Peng Xia
于 9月22日
今日最佳论文
·
Google
作者:
Peng Xia
,
Rujun Han
,
Zifeng Wang
,
Yanfei Chen
,
Yufan Zhang
,
Yoonho Lee
,
Chengsong Huang
,
Han Yu
,
Zhongying CuiZhu
,
Yifei Ming
,
Huaxiu Yao
,
Burak Gokturk
,
Tomas Pfister
,
Chen-Yu Lee
摘要
大型语言模型(LLM)智能体的能力在很大程度上由其工具链(harness)所放大,即围绕冻结骨干模型构建的提示词、控制流、工具、记忆和上下文管理机制。最近的方法通过迭代地提出并选择智能体工具链中各个组件的编辑方案,日益自动化这一过程,实际上在智能体系统层面确立了一种递归自改进(RSI)的形式。然而,这种递归进化可能会因记忆训练任务而过拟合,表现出巨大的分布内增益,但在分布外基准测试上这些增益会缩小甚至消失。我们引入了智能体工具链的正则化递归自改进(RRSI),通过将正则化原则融入工具链的自改进过程中,约束进化候选者的提出和选择。提议者采用时间退火预算,限制候选者可以捆绑的编辑数量,并基于进化历史鼓励探索未开发的轨迹。选择器配备了批评者和修剪器:批评者筛选针对特定基准的提案,而修剪器则移除那些变化过小、成本过高或不再有用的更改。这些约束共同 favor 可复用的智能体机制,而非针对特定基准的机制甚至噪声。在涵盖编码、智能体工作区和工程设计任务的八个基准测试中,RRSI 在其进化的拆分上最高获得 14.1 分的提升,在五个分布外基准测试上最高获得 4.7 分的提升,同时生成的工具链比未正则化的进化少消耗 30% 的策略令牌。代码可在 https://github.com/google-research/rrsi 获取,项目页面为 https://regularized-rsi.com/。
查看 arXiv 页面
查看 PDF
项目页面
GitHub
54
添加到合集
社区
richardxp888
论文作者
论文提交者
大约 19 小时前
我们发布了智能体工具链的正则化递归自改进!快来看看吧!
查看翻译
回复
编辑
预览
通过拖拽、粘贴或点击此处上传图片、音频和视频。
评论
· 注册或登录以发表评论
赞成票
152
+140
在你的智能体中获取此论文:
hf papers read 2609.24972
没有最新的 CLI?
引用此论文的模型
0
无链接此论文的模型
在模型的 README.md 中引用 arxiv.org/abs/2609.24972 以从此页面链接它。
引用此论文的数据集
0
无链接此论文的数据集
在数据集的 README.md 中引用 arxiv.org/abs/2609.24972 以从此页面链接它。
引用此论文的 Spaces
1
🛸
7Ali/https-globall-cloud-pages-dev
包含此论文的合集
1
paper2read
合集
189 项
·
更新于大约 7 小时前
·
7
Papers
arxiv:2609.24972
Copy markdown
RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
Published on Sep 21
·
Submitted by
Peng Xia
on Sep 22
·
Google
Authors:
Peng Xia
,
Rujun Han
,
Zifeng Wang
,
Yanfei Chen
,
Yufan Zhang
,
Yoonho Lee
,
Chengsong Huang
,
Han Yu
,
Zhongying CuiZhu
,
Yifei Ming
,
Huaxiu Yao
,
Burak Gokturk
,
Tomas Pfister
,
Chen-Yu Lee
Abstract
An LLM agent's capability is largely magnified by its harness, namely the prompts, control flow, tooling, memory, and context management surrounding the frozen backbone model. Recent methods increasingly automate this process by iteratively proposing and selecting component-wise edits of an agent harness, practically establishing a form of recursive self-improvement (RSI) at the agent-system level. However, such recursive evolution may overfit by memorizing the training tasks, showing large in-distribution gains that shrink or even vanish on out-of-distribution benchmarks. We introduce Regularized Recursive Self-Improvement of Agent Harnesses (RRSI), which incorporates the principles of regularizations into harness self-improvement by constraining the evolution candidate proposal and selection. The proposer operates with a temporally annealed budget, limiting how many edits a candidate can bundle, and it encourages unexplored trajectories based on evolution history. The selector is equipped with a critic and a pruner: the critic screens benchmark-specific proposals, while the pruner, removes changes that are too small, too expensive, or no longer useful. Together these constraints favor reusable agent mechanisms over benchmark-specific ones or even noises. Across eight benchmarks spanning coding, agentic workspace and engineering design tasks, RRSI gains up to 14.1 points on the split it evolves against and up to 4.7 points on the five out-of-distribution benchmarks, while producing a harness that runs on 30% fewer policy tokens than the unregularized evolution. Code is available at https://github.com/google-research/rrsi and project page is https://regularized-rsi.com/.
View arXiv page
View PDF
Project page
GitHub
54
Add to collection
Community
richardxp888
Paper author
Paper submitter
about 19 hours ago
We released the Regularized Recursive Self-Improvement of Agent Harnesses! Check it out!
See translation
Reply
Edit
Preview
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Comment
· Sign up or log in to comment
Upvote
152
+140
Get this paper in your agent:
hf papers read 2609.24972
Don't have the latest CLI?
Models citing this paper
0
No model linking this paper
Cite arxiv.org/abs/2609.24972 in a model README.md to link it from this page.
Datasets citing this paper
0
No dataset linking this paper
Cite arxiv.org/abs/2609.24972 in a dataset README.md to link it from this page.
Spaces citing this paper
1
🛸
7Ali/https-globall-cloud-pages-dev
Collections including this paper
1
paper2read
Collection
189 items
•
Updated about 7 hours ago
•
7
首次收录 · 2026-09-23 · 10.95 分