论文
arxiv:2609.35259
复制 Markdown
在线策略与离线策略学习?蒸馏动力学的系统研究
发布于 9月28日
·
提交者
Julianna Piskorz
于10月2日提交
·
剑桥大学
作者:
Julianna Piskorz
,
Antonin Berthon
,
Mihaela van der Schaar
摘要
在线策略学习被认为可以减少灾难性遗忘,产生更稀疏的参数更新,并提高泛化能力。然而,现有的监督微调与强化学习之间的比较同时变化了多个因素,使得 rollout 策略的贡献难以隔离。我们通过独立地改变 rollout 策略、词元级别的 KL 散度方向以及学习率,在 Llama3 和 Qwen2.5 模型系列以及涵盖科学、医学和算术领域的推理任务中,研究了受控的强到弱蒸馏设置下 rollout 策略的影响。我们的分析揭示了一个复杂的蒸馏动力学图景,其中 rollout 策略并不一定起核心作用。相反,词元级别的 KL 散度方向更清晰地塑造了任务表现和输出覆盖范围,而学习率则控制着遗忘和更新稀疏性。对 KL 梯度的分析以及沿连续的学生-教师 rollout 策略谱系的实验解释了这一模式:前向 KL 对 rollout 策略具有显著的鲁棒性,尽管 rollout 策略发生变化,其表现依然稳定且强劲;而后向 KL 则敏感得多,并偏好由学生生成的 rollout。尽管如此,在线策略数据在两种 KL 散度方向下均能改善对 Countdown 算术任务更困难变体的泛化能力,尽管这种优势在随后的 RLVR(强化学习价值回归)后并不总能持续存在。我们的更广泛结论在移除梯度裁剪、使用采样 KL 估计器以及训练需要更长推理链的任务时依然稳健。总体而言,我们的结果挑战了在线策略 rollout 本质上更优越的观点,并表明其价值关键取决于目标函数、评估设置和优化超参数。
查看 arXiv 页面
查看 PDF
添加到合集
社区
jpiskorz
论文提交者
大约7小时前
好奇听听你们对于在强到弱蒸馏背景下,在线与离线策略学习之间差异的看法和评论!
查看翻译
回复
编辑
预览
通过拖拽、粘贴或点击此处上传图片、音频和视频。
评论
· 注册或登录以发表评论
赞成
114
+102
在你的智能体中获取此论文:
hf papers read 2609.35259
没有最新的 CLI?
引用此论文的模型
0
无链接此论文的模型
在模型的 README.md 中引用 arxiv.org/abs/2609.35259 以从本页链接它。
引用此论文的数据集
0
无链接此论文的数据集
在数据集的 README.md 中引用 arxiv.org/abs/2609.35259 以从本页链接它。
引用此论文的 Spaces
0
无链接此论文的 Space
在 Space 的 README.md 中引用 arxiv.org/abs/2609.35259 以从本页链接它。
包含此论文的合集
0
无包含此论文的合集
将此论文添加到合集中以从本页链接它。
Papers
arxiv:2609.35259
Copy markdown
On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics
Published on Sep 28
·
Submitted by
Julianna Piskorz
on Oct 2
·
University of Cambridge
Authors:
Julianna Piskorz
,
Antonin Berthon
,
Mihaela van der Schaar
Abstract
On-policy learning has been argued to reduce catastrophic forgetting, produce sparser parameter updates, and improve generalisation. However, existing comparisons between supervised fine-tuning and reinforcement learning vary many factors simultaneously, making the contribution of rollout policy difficult to isolate. We study the effect of rollout policy in a controlled strong-to-weak distillation setting, by independently varying rollout policy, token-level KL direction, and learning rate across the Llama3 and Qwen2.5 model families and reasoning tasks spanning scientific, medical, and arithmetic domains. Our analysis reveals a nuanced picture of distillation dynamics in which rollout policy does not necessarily play a central role. Instead, token-level KL direction more clearly shapes task performance and output coverage, while learning rate governs forgetting and update sparsity. Analysis of KL gradients and experiments along a continuous student-teacher rollout-policy spectrum explain this pattern: forward KL is remarkably robust to rollout policy, with its performance stable and strong despite changes to the rollout policy, whereas reverse KL is substantially more sensitive and favours student-generated rollouts. On-policy data nevertheless improves generalisation to harder variants of the Countdown arithmetic task under both KL directions, although this advantage does not reliably persist after subsequent RLVR. Our broader conclusions remain robust to removing gradient clipping, using sampled KL estimators, and training on tasks requiring longer reasoning chains. Overall, our results challenge the view that on-policy rollouts are inherently preferable and show that their value depends critically on the objective, evaluation setting, and optimisation hyperparameters.
View arXiv page
View PDF
Add to collection
Community
jpiskorz
Paper submitter
about 7 hours ago
Curious to hear your thoughts and comments about the differences between on- and off-policy learning in the context of strong-to-weak distillation!
See translation
Reply
Edit
Preview
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Comment
· Sign up or log in to comment
Upvote
114
+102
Get this paper in your agent:
hf papers read 2609.35259
Don't have the latest CLI?
Models citing this paper
0
No model linking this paper
Cite arxiv.org/abs/2609.35259 in a model README.md to link it from this page.
Datasets citing this paper
0
No dataset linking this paper
Cite arxiv.org/abs/2609.35259 in a dataset README.md to link it from this page.
Spaces citing this paper
0
No Space linking this paper
Cite arxiv.org/abs/2609.35259 in a Space README.md to link it from this page.
Collections including this paper
0
No Collection including this paper
Add this paper to a collection to link it from this page.
首次收录 · 2026-10-03 · 10.95 分