论文
arxiv:2609.22068
复制 Markdown
CodeMidas:从代码本身扩展智能体编码强化学习环境
发布于 9月18日
·
提交者
李磊
于 9月21日
·
小米 MiMo
作者:
叶博文
,
李磊
,
李诗成
,
岳子豪
,
张凌浩
,
吕航龙
,
刘元鑫
,
马文翰
,
田浩
,
李让
,
董金浩
,
赵一凯
,
邓祥伟
,
张海林
,
赵亮
,
刘琦
,
孔令鹏
,
杨通
,
罗富丽
摘要
通过强化学习(RL)训练具备能力的编码智能体需要多样化的任务以及可靠的验证器。开源代码库提供了丰富的此类任务来源,而现有方法通常依赖于问题(issues)和提交(commits)等开发工件,限制了可提取的任务范围。为了更好地扩展强化学习环境,我们提出了 CodeMidas,这是一种智能体管道,它将现有代码库中已实现的功能转化为可执行的强化学习环境,仅以源代码作为特定任务的输入。CodeMidas 为环境构建的每个阶段分配智能体计算资源:智能体探索已实现的功能以制定行为规范,基于原始代码的执行构建测试,并通过执行检查和重复解决方案 rollout 来验证和过滤候选任务。生成的数据集包含来自 3,185 个开源代码库的 5,545 个训练任务,涵盖 23 种编程语言和 15 个技术领域。在 GRPO 的帮助下,在这些任务上训练 MiMo-V2.5 提升了所有五个多样化基准测试的性能,包括问题修复(DeepSWE + 11.7%)、全程序构建(ProgramBench + 17%)以及终端工作(Terminal-Bench v2.1 + 8.5%)。消融实验表明,增加高质量训练任务的数量可以提升性能。轨迹分析显示,经过强化学习的智能体表现出更好的行为,如增加代码库探索频率和更多样化的自我验证。这些结果确立了源代码作为构建强化学习环境的可扩展基础,从而提升编码智能体在多样化软件任务中的表现。
查看 arXiv 页面
查看 PDF
项目页面
添加到合集
社区
tobiaslee
论文提交者
大约 20 小时前
我们提出了 CodeMIDAS,旨在从代码本身扩展智能体编码强化学习环境。
https://mimo.xiaomi.com/rl/
查看翻译
🔥
1
+
回复
researchstudio-bot
大约 14 小时前
这是 ResearchStudio 团队发出的自动消息。
我们为这篇论文创建了一个互动的 ResearchStudio Reel。它包括视觉海报、视频和博客,所有文件均可下载且格式可编辑。
打开 ResearchStudio Reel →
从 Hugging Face 下载所有文件
如果您觉得这个 Reel 有帮助,请给此评论点赞!
想探索或为更多论文创建 Reels?请访问 ResearchStudio 演示。
查看翻译
回复
编辑
预览
通过拖拽、粘贴或点击此处上传图片、音频和视频。
评论
· 注册或登录以发表评论
赞同
86
+74
在您的智能体中获取此论文:
hf papers read 2609.22068
没有最新的 CLI?
引用此论文的模型
0
无链接此论文的模型
在模型的 README.md 中引用 arxiv.org/abs/2609.22068 以从此页面链接它。
引用此论文的数据集
0
无链接此论文的数据集
在数据集的 README.md 中引用 arxiv.org/abs/2609.22068 以从此页面链接它。
引用此论文的 Spaces
0
无链接此论文的 Space
在 Space 的 README.md 中引用 arxiv.org/abs/2609.22068 以从此页面链接它。
包含此论文的合集
0
无包含此论文的合集
将此论文添加到合集中以从此页面链接它。
Papers
arxiv:2609.22068
Copy markdown
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself
Published on Sep 18
·
Submitted by
Lei Li
on Sep 21
·
Xiaomi MiMo
Authors:
Bowen Ye
,
Lei Li
,
Shicheng Li
,
Zihao Yue
,
Linghao Zhang
,
Hanglong Lv
,
Yuanxin Liu
,
Wenhan Ma
,
Hao Tian
,
Rang Li
,
Jinhao Dong
,
Yikai Zhao
,
Xiangwei Deng
,
Hailin Zhang
,
Liang Zhao
,
Qi Liu
,
Lingpeng Kong
,
Tong Yang
,
Fuli Luo
Abstract
Training capable coding agents via reinforcement learning (RL) requires diverse tasks with reliable verifiers. Open-source codebases offer a rich source of such tasks, while existing methods typically rely on development artifacts such as issues and commits, limiting the range of tasks that can be extracted. To better scale RL environments, we present CodeMidas, an agentic pipeline that turns implemented functionality in existing codebases into executable RL environments using source code as its only task-specific input. CodeMidas allocates agentic compute to every stage of environment construction: agents explore implemented functionality to formulate behavioral specifications, construct tests grounded in execution of the original code, and validate and filter candidate tasks through execution checks and repeated solution rollouts. The resulting dataset has 5,545 training tasks from 3,185 open-source codebases spanning 23 programming languages and 15 technical domains. Training MiMo-V2.5 on these tasks with GRPO improves performance on all five diverse benchmarks, covering issue repair (DeepSWE + 11.7%), whole-program construction (ProgramBench +17%), and terminal work (Terminal-Bench v2.1 +8.5%). Ablations show that increasing the number of high-quality training tasks improves performance. Trajectory analysis shows the RL-trained agent demonstrates better behaviors like increasing codebase exploration and more diverse self-verification. These results establish source code as a scalable foundation for constructing RL environments that improve coding agents across diverse software tasks.
View arXiv page
View PDF
Project page
Add to collection
Community
tobiaslee
Paper submitter
about 20 hours ago
We present CodeMIDAS towards scaling agentic coding RL environments from code itself.
https://mimo.xiaomi.com/rl/
See translation
🔥
1
+
Reply
researchstudio-bot
about 14 hours ago
This is an automated message from the ResearchStudio team.
We created an interactive ResearchStudio Reel for this paper. It includes a visual poster, a video, and a blog, all available for download in editable formats.
Open the ResearchStudio Reel →
Download all files from Hugging Face
Please give this comment a thumbs up if you find the Reel helpful!
Want to explore or create Reels for more papers? Visit the ResearchStudio demo.
See translation
Reply
Edit
Preview
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Comment
· Sign up or log in to comment
Upvote
86
+74
Get this paper in your agent:
hf papers read 2609.22068
Don't have the latest CLI?
Models citing this paper
0
No model linking this paper
Cite arxiv.org/abs/2609.22068 in a model README.md to link it from this page.
Datasets citing this paper
0
No dataset linking this paper
Cite arxiv.org/abs/2609.22068 in a dataset README.md to link it from this page.
Spaces citing this paper
0
No Space linking this paper
Cite arxiv.org/abs/2609.22068 in a Space README.md to link it from this page.
Collections including this paper
0
No Collection including this paper
Add this paper to a collection to link it from this page.
首次收录 · 2026-09-22 · 10.95 分