论文
arxiv:2609.35347
复制 Markdown
超越教师分配:领域归一化的多教师在线策略蒸馏
发布于 9月28日
·
提交者
XinLi
于 9月29日
·
南洋理工大学
作者:
Xin Li
,
Hao Jiang
,
Xin Gao
,
Annan Wang
,
Yuchen Xie
,
Jinghao Guo
,
Xingwei Qu
,
Yichi Zhang
,
Chau Yuen
摘要
强化学习可以将一个语言模型转化为多个专家,每个专家都精通于一项特定技能,如数学、编程或遵循指令,但用户需要一个具备所有这些技能的模型。多教师在线策略蒸馏(MOPD)通过让专家指导一名学生来合并这些能力:学生回答每个提示,而该提示所属领域的专家会对每一个token提供反馈。这种路由机制决定了哪位专家进行教学,但未决定其反馈对共享学生的影响强度。在三个不同规模的 Qwen3.5 模型中,我们发现 MOPD 的学生并未超越由最佳单一专家指导的学生,且几乎未获得数学专家的优势。反馈是不平衡的:遵循指令的反馈分布范围是数学反馈的数倍,并主导了学生的更新过程。我们提出了领域归一化的 MOPD(DN-MOPD),它保留了路由机制,并根据测量到的分布范围对每个领域的反馈进行重新缩放。在六个公开基准测试中,DN-MOPD 在所有规模、三种随机种子以及两种答案长度限制下,均提升了平均分数,并恢复了大部分丢失的数学能力增益。使用固定领域权重的对照实验表明,这种增益主要来自于降低遵循指令反馈的影响,而非单纯提高数学反馈;且接近 DN-MOPD 测量值的固定权重表现相当。因此,结合专家不仅需要决定由谁进行教学,还需要决定其反馈的重要性程度。
查看 arXiv 页面
查看 PDF
项目主页
GitHub
10
添加到合集
社区
XINLI1997
论文作者
论文提交者
约 17 小时前
项目主页:https://lixin.ai/DN-MOPD/ 。代码:https://github.com/LiXin97/DN-MOPD
查看翻译
回复
编辑
预览
通过拖拽、粘贴到文本输入框或点击此处来上传图片、音频和视频。
评论
· 注册或登录以发表评论
赞同
138
+126
在您的智能体中获取此论文:
hf papers read 2609.35347
没有最新的 CLI?
引用此论文的模型
0
无链接此论文的模型
在模型的 README.md 中引用 arxiv.org/abs/2609.35347 以从此页面链接它。
引用此论文的数据集
1
XINLI1997/DN-MOPD-Data
查看器
•
更新于约 17 小时前
•
26.7k
•
56
引用此论文的 Spaces
0
无链接此论文的 Space
在 Space 的 README.md 中引用 arxiv.org/abs/2609.35347 以从此页面链接它。
包含此论文的合集
1
DN-MOPD
合集
超越教师分配:领域归一化的多教师在线策略蒸馏。数据和模型。
•
2 项内容
•
更新于约 17 小时前
Papers
arxiv:2609.35347
Copy markdown
Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation
Published on Sep 28
·
Submitted by
XinLi
on Sep 29
·
Nanyang Technological University
Authors:
Xin Li
,
Hao Jiang
,
Xin Gao
,
Annan Wang
,
Yuchen Xie
,
Jinghao Guo
,
Xingwei Qu
,
Yichi Zhang
,
Chau Yuen
Abstract
Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions, but users need one model with all of these skills. Multi-teacher on-policy distillation (MOPD) merges them by letting the specialists teach one student: the student answers each prompt, and the specialist for that prompt's domain gives feedback on every token. This routing decides which specialist teaches, but not how strongly its feedback moves the shared student. In Qwen3.5 models at three sizes, we find that MOPD's student does not beat one taught by the best single specialist and gains little of the mathematics specialist's advantage. The feedback is unbalanced: instruction-following feedback is several times more spread out than mathematics feedback and dominates the student's updates. We propose Domain-Normalized MOPD (DN-MOPD), which keeps the routing and rescales each domain's feedback by its measured spread. On six public benchmarks, DN-MOPD improves the average score over MOPD at every size, across three random seeds and under two answer-length limits, and recovers most of the lost mathematics gain. Controls with fixed domain weights show that the gain comes mainly from turning down instruction-following feedback rather than turning up mathematics alone, and that fixed weights close to those DN-MOPD measures perform comparably. Combining specialists therefore requires deciding not only which one teaches, but also how strongly its feedback counts.
View arXiv page
View PDF
Project page
GitHub
10
Add to collection
Community
XINLI1997
Paper author
Paper submitter
about 17 hours ago
Project page: https://lixin.ai/DN-MOPD/ . Code: https://github.com/LiXin97/DN-MOPD
See translation
Reply
Edit
Preview
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Comment
· Sign up or log in to comment
Upvote
138
+126
Get this paper in your agent:
hf papers read 2609.35347
Don't have the latest CLI?
Models citing this paper
0
No model linking this paper
Cite arxiv.org/abs/2609.35347 in a model README.md to link it from this page.
Datasets citing this paper
1
XINLI1997/DN-MOPD-Data
Viewer
•
Updated about 17 hours ago
•
26.7k
•
56
Spaces citing this paper
0
No Space linking this paper
Cite arxiv.org/abs/2609.35347 in a Space README.md to link it from this page.
Collections including this paper
1
DN-MOPD
Collection
Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation. Data and models.
•
2 items
•
Updated about 17 hours ago
首次收录 · 2026-09-30 · 10.95 分