论文
arxiv:2609.29845
复制 Markdown
你的 Transformer 能同时持有两个想法:大语言模型中线性叠加的证据
发布于 9月24日
·
提交者
Anton Razzhigaev
于 9月25日
作者:
Pavel Tikhonov
,
Anton Korznikov
,
Matvey Mikhalchuk
,
Nikita Dragunov
,
Temurbek Rahmatullaev
,
Polina Druzhinina
,
Anton Razzhigaev
,
Ivan Oseledets
,
Elena Tutubalina
摘要
尽管大语言模型(LLMs)依赖于高度非线性的组件,但在这项工作中,我们证明了它们表现出基本的线性特性:当来自不同文本流的输入被线性组合时,模型会输出各个单独下一个 token 分布的叠加。我们将此称为“叠加线性假设”(Superposition Linearity Hypothesis)。我们提供了证据表明,叠加是 Transformer 架构的一种固有属性,而非训练产生的衍生结果;事实上,我们观察到随着预训练的推进,这种特性往往会减弱。然而,我们通过轻量级微调证明了可以大幅恢复这种线性,显著降低了预测的下一个 token 分布与各个单独下一个 token 分布平均值之间的差异。最后,我们引入了一种引导解码过程,用于解纠缠叠加的输出,从而能够从单次前向传播中同时生成两个连贯的后续内容。
查看 arXiv 页面
查看 PDF
添加到合集
社区
razzant
论文作者
论文提交者
大约 9 小时前
你的 LLM 能同时持有两个想法:当你将两个不相关文本的嵌入向量取平均时,模型会同时预测这两个流的下一个 token。这种叠加源于架构本身,因为预训练会侵蚀它,而轻量级微调可以恢复它。我们展示了如何再次将其分离,并从单次前向传播中解码出两个后续内容。
查看翻译
❤️
3
+
回复
编辑
预览
通过拖拽、粘贴或点击此处上传图片、音频和视频到文本输入框。
评论
· 注册或登录以发表评论
赞同
55
+43
在您的智能体中获取此论文:
hf papers read 2609.29845
没有最新的 CLI?
引用此论文的模型
0
无链接此论文的模型
在模型的 README.md 中引用 arxiv.org/abs/2609.29845 以从此页面链接它。
引用此论文的数据集
0
无链接此论文的数据集
在数据集的 README.md 中引用 arxiv.org/abs/2609.29845 以从此页面链接它。
引用此论文的 Spaces
0
无链接此论文的 Space
在 Space 的 README.md 中引用 arxiv.org/abs/2609.29845 以从此页面链接它。
包含此论文的合集
1
WTF GENIUS PAPERS
合集
让我稍微更欣赏自己的专业和生活的论文。obs=观察,innov=创新。大多数论文旨在改进小型模型。
•
409 项
•
更新于大约 6 小时前
•
87
Papers
arxiv:2609.29845
Copy markdown
Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs
Published on Sep 24
·
Submitted by
Anton Razzhigaev
on Sep 25
Authors:
Pavel Tikhonov
,
Anton Korznikov
,
Matvey Mikhalchuk
,
Nikita Dragunov
,
Temurbek Rahmatullaev
,
Polina Druzhinina
,
Anton Razzhigaev
,
Ivan Oseledets
,
Elena Tutubalina
Abstract
While Large Language Models (LLMs) rely on highly non-linear components, in this work we demonstrate that they exhibit fundamental linearity: when inputs from distinct text streams are linearly combined, the model outputs a superposition of the individual next-token distributions. We term this the Superposition Linearity Hypothesis. We provide evidence that superposition is an intrinsic property of the Transformer architecture rather than an emergent consequence of training; in fact, we observe that it tends to diminish as pretraining progresses. However, we demonstrate that linearity can be substantially restored through lightweight fine-tuning, significantly reducing the divergence between the predicted next-token distribution and the average of the individual next-token distributions. Finally, we introduce a guided decoding procedure that disentangles superposed outputs, enabling the simultaneous generation of two coherent continuations from a single forward pass.
View arXiv page
View PDF
Add to collection
Community
razzant
Paper author
Paper submitter
about 9 hours ago
Your LLM can hold two thoughts at once: when you average the embeddings of two unrelated texts, the model predicts the next token for both streams at the same time. This superposition comes from the architecture itself, since pretraining erodes it and a light fine-tune restores it. We show how to separate it again and decode two continuations from a single forward pass.
See translation
❤️
3
+
Reply
Edit
Preview
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Comment
· Sign up or log in to comment
Upvote
55
+43
Get this paper in your agent:
hf papers read 2609.29845
Don't have the latest CLI?
Models citing this paper
0
No model linking this paper
Cite arxiv.org/abs/2609.29845 in a model README.md to link it from this page.
Datasets citing this paper
0
No dataset linking this paper
Cite arxiv.org/abs/2609.29845 in a dataset README.md to link it from this page.
Spaces citing this paper
0
No Space linking this paper
Cite arxiv.org/abs/2609.29845 in a Space README.md to link it from this page.
Collections including this paper
1
WTF GENIUS PAPERS
Collection
Papers that made me appreciate my major and my life a little more. obs=Observation, innov=Innovation. Most papers are abt improving tiny models.
•
409 items
•
Updated about 6 hours ago
•
87
首次收录 · 2026-09-26 · 10.95 分