论文
arxiv:2609.29233
复制 Markdown
训练后行为在无关决策中留下“行为阴影”
发布于 9月24日
·
提交者
曾元浩
于 9月29日
·
北京大学
作者:
张子阳
,
景宇斌
,
曾元浩
,
李裕瑶
,
王浩然
,
龚一辰
摘要
我们发现,语言模型可以通过与任务无关的文本转移能力。后训练通常使用特定任务数据来改进语言模型。关于潜意识学习的前期研究表明,有关这些更新的信息可以通过无关生成传递,但主要侧重于使用大量教师输出得出的特征或偏好。我们引入了无任务主动蒸馏(Active Taskless Distillation, ATD),该方法仅通过每个提示中来自教师的单个词即可实现能力转移。ATD 通过后训练留下的“行为阴影”探测能力:选择那些教师和学生模型共享的公共祖先对两个普通词语几乎无偏好的提示。从该祖先初始化的学生仅从生成的提示-词对中学习,无需目标任务示例、教师 logits 或教师参数。在 Qwen2.5-1.5B 的主要编码实验中,5,664 个样本使得 HumanEval+ 上的准确率比精确干扰匹配的控制组提高了 5.34 个百分点。其他实验表明,科学常识、常识推理和阅读理解等能力在不同模型代际、尺寸和家族中也存在转移。功能分析表明,所学能力是可组合的,且其强度与教师的更新强度呈正相关。
查看 arXiv 页面
查看 PDF
GitHub
13
添加到收藏
社区
评论
论文提交者
大约 17 小时前
一个模型能否在不展示任何代码的情况下教另一个模型编程?
令人惊讶的是,可以。一个经过编码训练的教师仅用单个词回答数千个无关提示。一个仅基于这些答案进行训练的学生在编程方面变得更强——尽管它从未见过代码、任务示例或教师 logits。仅通过 5,664 个单字答案,该学生在 HumanEval+ 上比干扰匹配的控制组提高了 +5.34 个百分点。
为什么?后训练留下了分布式的“行为阴影”:即使在无关语境中,普通词语之间也存在微小的偏好偏移。我们的方法——无任务主动蒸馏(ATD)——针对基础模型几乎未决的决策点,即微小偏移揭示更新方向的位置。前期研究表明偏好可以通过这种方式传递;据我们所知,这是首次证明一种能力可以如此转移。
该效应在科学、常识、阅读理解、模型尺寸和家族中均成立。结论是:所学能力会溢出到其训练任务之外——数千个看似无意义的行为变化共同携带了后训练的丰富痕迹。
查看翻译
👍
2
+
回复
编辑
预览
通过拖拽、粘贴或点击此处上传图片、音频和视频。
评论
· 注册或登录以发表评论
赞同
249
+237
将此论文获取到你的代理中:
hf papers read 2609.29233
没有最新的 CLI?
引用此论文的模型
0
无链接此论文的模型
在模型 README.md 中引用 arxiv.org/abs/2609.29233 以从此页链接它。
引用此论文的数据集
0
无链接此论文的数据集
在数据集 README.md 中引用 arxiv.org/abs/2609.29233 以从此页链接它。
引用此论文的 Spaces
0
无链接此论文的 Space
在 Space README.md 中引用 arxiv.org/abs/2609.29233 以从此页链接它。
包含此论文的集合
1
微调
集合
43 项
•
更新于大约 6 小时前
•
4
Papers
arxiv:2609.29233
Copy markdown
Post-Training Leaves Behavioral Shadows on Unrelated Decisions
Published on Sep 24
·
Submitted by
Zeng Yuanhao
on Sep 29
·
Peking University
Authors:
Ziyang Zhang
,
Yubin Jing
,
Yuanhao Zeng
,
Yuyao Li
,
Haofan Wang
,
Yichen Gong
Abstract
We find that language models can transfer capabilities through task-unrelated text. Post-training typically improves language models using task-specific data. Prior work on subliminal learning shows that information about these updates can pass through unrelated generations, but has largely focused on traits or preferences using extensive teacher outputs. We introduce Active Taskless Distillation (ATD), which achieves capability transfer using only a single word from the teacher per prompt. ATD probes the behavioral shadow of post-training by selecting prompts where the teacher and student's shared public ancestor is nearly indifferent between two ordinary words. A student initialized from this ancestor learns solely from the resulting prompt-word pairs, without target-task examples, teacher logits, or teacher parameters. In the primary coding experiment with Qwen2.5-1.5B, 5,664nses yield a 5.34 pp gain on HumanEval+ over an exact nuisance-matched control thadisrupts prompt-resperiments showtransfer in scientific knowledge, commonsense reasoning, and reading comprehensins across additional model generations, sizes, and families. Functional analyses show that the learned sid composable, andthat its strength tracks the teacher's update strength.
View arXiv page
View PDF
GitHub
13
Add to collection
Community
Lines
Paper submitter
about 17 hours ago
Can a model teach another model to code — without ever showing it code?
Surprisingly, yes. A coding-trained teacher answers thousands of unrelated prompts with just one word each. A student trained only on those answers gets better at coding — despite never seeing code, task examples, or teacher logits. With just 5,664 single-word answers, the student gains +5.34pp on HumanEval+ over a nuisance-matched control.
Why? Post-training leaves a distributed "behavioral shadow": tiny preference shifts between ordinary words, even in irrelevant contexts. Our method, Active Taskless Distillation (ATD), targets decisions where the base model is nearly undecided — where small shifts reveal the direction of the update. Prior work showed that preferences can be transmitted this way; to our knowledge, this is the first time a capability is.
The effect holds across science, commonsense, reading comprehension, model sizes, and families. The punchline: learned capabilities leak far beyond their training task — thousands of seemingly meaningless behavioral changes collectively carry a rich trace of post-training.
See translation
👍
2
+
Reply
Edit
Preview
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Comment
· Sign up or log in to comment
Upvote
249
+237
Get this paper in your agent:
hf papers read 2609.29233
Don't have the latest CLI?
Models citing this paper
0
No model linking this paper
Cite arxiv.org/abs/2609.29233 in a model README.md to link it from this page.
Datasets citing this paper
0
No dataset linking this paper
Cite arxiv.org/abs/2609.29233 in a dataset README.md to link it from this page.
Spaces citing this paper
0
No Space linking this paper
Cite arxiv.org/abs/2609.29233 in a Space README.md to link it from this page.
Collections including this paper
1
Fine tuning
Collection
43 items
•
Updated about 6 hours ago
•
4
首次收录 · 2026-09-30 · 10.95 分