控制框架决定代理在修改文件前是否读取了正确的文件,是否能够从错误中恢复,以及是否清晰地交付结果。根据一篇新的研究论文,近期代理领域的许多进展来自于对控制框架的工作,而非新模型的出现。
直到最近,这项工作都是人工完成的。人们审查失败的运行记录并手动修补控制框架。较新的方法通过让语言模型基于测试任务的反馈反复重写控制框架本身,从而自动化这一循环。
研究人员称这是一种实用的递归自我改进形式。该系统产生用于优化控制框架的反馈,而该控制框架反过来又控制系统自身的行为。
自我优化导致代理记忆其测试任务
论文表明,这种自我优化伴随着一个隐患。由于代理持续处理同一组有限的测试任务,它最终会记住这些任务。其在训练任务上的得分上升,而在新的、未见过的任务上的收益则缩小甚至完全消失。
研究人员表示,这种情况以多种方式发生。搜索过程记住了仅适用于特定基准的模式,偏爱那些纯粹因运气好而得分高的候选方案,并堆积不必要的复杂性,从而在不提升代理能力的情况下提高测试分数。
其他方法主要在训练任务上取得改进,但很少能迁移到未见过的基准上。RRSI 在三个领域均提升了这些基准的得分。| 图片来源:Google
不断缩小的编辑预算和严格的评审者保持控制框架的通用性
RRSI(代理控制框架的正则递归自我改进)在优化循环的两端发挥作用,同时保持控制框架完全可编辑。当系统提出新更改时,预算限制了候选方案一次可以捆绑多少个独立编辑。
该预算随时间推移而缩小。早期轮次允许较大的重写,而后期轮次仅允许那些能清晰追溯到结果的微小更改。系统还会记录早期的尝试,以免不断追逐同样的失败想法。当进展停滞时,它会故意对尚未触及的控制框架部分进行实验。
RRSI 在两个环节限制自我优化:提出新更改时,以及决定哪些更改成为控制框架永久组成部分时。| 图片来源:Google
在选择更改时,评审者会审查每个提案,并剔除任何硬编码任务名称、解决方案或其他基准特定技巧的提案。另一条规则规定,只有当计算成本带来可衡量的性能提升时,才接受更高的计算成本。不再有帮助的组件会被移除。
放弃训练收益在新任务上获得回报
研究人员在涵盖编码、代理办公工作和工程设计领域的八个基准上测试了 RRSI。底层模型 Claude Opus 4.8 在整个过程中保持冻结状态。团队将 RRSI 与未修改的基线控制框架以及四种最近的优化方法进行了比较。
根据论文,RRSI 在其训练任务上的得分最高提升 14.1 分,在五个从未见过的基准上最高提升 4.7 分。与未正则化的版本相比,它在运行时使用的令牌减少了约 30%。在所有未见过的基准上,整体性能从未低于基线水平,而这通常是控制框架记忆其任务时会发生的情况。
RRSI 控制框架也在训练集之外的每个任务上有所改进,在 JobBench 上的最大增益为 4.7 分。| 图片来源:Google
每种方法在训练任务上都表现良好,但在新任务上的结果发生了逆转。两种方法的得分甚至低于基线控制框架。RRSI 在所有变体中实现了最小的训练收益,并且是唯一在未见过任务上显著高于基线的方法。这些护栏旨在产生完全这种权衡。
在优化的控制框架中,RRSI 需要的令牌和步骤最少,并在新任务上表现最佳,尽管未修改的基线控制框架更加精简。| 图片来源:Google
在一款模型上优化的工具链也能帮助较弱的模型
使用 Gemini 3.5 Flash 优化的编码工具链,在未进行任何修改的情况下,将性能弱得多的 Gemini 3.1 Flash Lite 的准确率提高了 14.6 个百分点,从 11.2 提升至该数值。该系统发现的机制并不依赖于用于发现它们的模型的能力。
作者指出,他们的研究仅涵盖围绕冻结模型构建的工具链,并未涉及模型权重发生变化的情况。
他们得出结论,只有当重复反馈转化为持久性改变时,自我改进才能使 AI 代理可靠地提升能力。代码已在 GitHub 上提供。
手动设计的工具链往往无法泛化到新任务中,ARC-AGI-3 上的测试已证实了这一点。使用专为特定目的构建的工具链,Opus 4.6 在熟悉的环境中得分达到 97.1%,而在陌生环境中得分为 0%。英伟达(Nvidia)最近展示了 SoL-Pi ,这是一种相关方法,其中研究代理会自动重建编码代理的工具链。它在性能没有明显下降的情况下,将令牌使用量减少了高达 49%。
在此前不久,谷歌让代理“梦回”过去的搜索运行过程,以改进其搜索策略。这项工作同样未改变模型本身。
The harness decides whether an agent reads the right file before changing it, whether it recovers from a mistake, and whether it delivers its results cleanly. According to a new research paper, much of the recent progress in agents comes from work on the harness, not from new models.
Until recently, this was done by hand. People reviewed failed runs and patched the harness manually. Newer methods automate the loop by having a language model rewrite the harness itself, again and again, based on feedback from the test tasks.
The researchers call this a practical form of recursive self-improvement. The system produces feedback that it uses to optimize the harness, which in turn controls the system's own behavior.
Self-optimization leads agents to memorize their test tasks
The paper shows that this self-optimization comes with a catch. Because the agent keeps working on the same limited set of test tasks, it ends up memorizing them. Its scores on the training tasks go up, while gains on new, unseen tasks shrink or disappear entirely.
The researchers say this happens in several ways. The search memorizes patterns that only fit one particular benchmark, favors candidates that score well purely by chance, and piles on unnecessary complexity that raises the test score without making the agent any better.
Other methods mostly improve on the training tasks, but little of that carries over to unseen benchmarks. RRSI raises scores there in all three domains. | Image: Google
Shrinking edit budgets and a strict critic keep the harness general
RRSI (Regularized Recursive Self-Improvement of Agent Harnesses) works on both ends of the optimization loop while leaving the harness fully editable. When the system proposes new changes, a budget caps how many independent edits a candidate can bundle at once.
That budget shrinks over time. Early rounds allow larger rewrites, while later rounds only permit small changes that can be clearly traced to a result. The system also keeps track of earlier attempts so it doesn't keep chasing the same failed ideas. When progress stalls, it deliberately experiments with parts of the harness it hasn't touched yet.
RRSI reins in self-optimization at two points, when proposing new changes and when deciding which of them become a permanent part of the harness. | Image: Google
When it comes to picking changes, a critic reviews every proposal and throws out any that hardcode task names, solutions, or other benchmark-specific tricks. Another rule only accepts higher compute costs if they come with a measurable performance gain. Components that no longer help get removed.
Giving up training gains pays off on new tasks
The researchers tested RRSI on eight benchmarks spanning coding, agentic office work, and engineering design. The underlying model, Claude Opus 4.8 , stayed frozen throughout. The team compared RRSI with the unmodified baseline harness and four recent optimization methods.
According to the paper, RRSI gains up to 14.1 points on the tasks it was trained on and up to 4.7 points on five benchmarks it never saw. It also uses about 30 percent fewer tokens at runtime than the unregularized version. Overall performance never fell below the baseline on any of the unseen benchmarks, which typically happens with a harness that has memorized its tasks.
The RRSI harness also improves on every task outside the training set, with the biggest gain of 4.7 points on JobBench. | Image: Google
Every method did well on the training tasks, but the results flipped on new ones. Two methods even ended up below the baseline harness. RRSI posted the smallest training gain of all the variants and was the only method that landed well above the baseline on unseen tasks. The guardrails are meant to produce exactly this tradeoff.
Among the optimized harnesses, RRSI needs the fewest tokens and steps and performs best on new tasks, though the unmodified baseline harness is even leaner. | Image: Google
Harnesses optimized on one model also help weaker ones
A coding harness optimized with Gemini 3.5 Flash raised the accuracy of the much weaker Gemini 3.1 Flash Lite from 11.2 to 14.6 points without any modifications. The mechanisms the system found don't depend on the capability of the model used to discover them.
The authors note that their study only covers harnesses built around frozen models and doesn't address cases where the model weights change.
They conclude that self-improvement only makes AI agents reliably more capable when repeated feedback gets turned into lasting changes. The code is available on GitHub .
Manually designed harnesses often don't generalize to new tasks, as tests on ARC-AGI-3 have shown. With a purpose-built harness, Opus 4.6 scored 97.1 percent in a familiar environment and 0 percent in an unfamiliar one. Nvidia recently presented SoL-Pi , a related method in which a research agent automatically rebuilds the harness of coding agents. It cuts token use by up to 49 percent without a noticeable drop in performance.
Shortly before that, Google had agents "dream" about past search runs to improve their search strategy. That work also leaves the model itself unchanged.
首次收录 · 2026-10-05 · 10.48 分