论文
arxiv:2609.37533
复制 Markdown
E-MoE:用于非因式分解扩散语言模型的增强型混合专家模型
发布于 9月29日
·
提交者
Arseny Ivanov
于10月2日
作者:
Arseny Ivanov
,
Alexander Kolesov
,
Alexander Korotin
,
Ivan Oseledets
,
Mikhail Goncharov
摘要
掩码扩散模型(MDMs)通过在每个去噪步骤中逐步解除对多个标记的掩码来生成序列,但其逆过程通常按位置进行因式分解,这在扩散速度优势相对于自回归解码最为重要的少步生成场景中限制了样本质量。最近的一系列工作引入了一种连续高斯潜在变量,作为变分自编码器进行训练,以捕捉跨位置的关联性,但此类方法容易出现后验坍缩,即潜在变量被无声地忽略。我们提出了增强型混合专家模型(E-MoE),它将逆过程构建为基于离散共享潜在变量的因式分解分布的混合体,该离散共享潜在变量由混合专家(MoE)主干网的专家路由决策给出,且不会增加相对于因式分解基线的活跃参数数量。在合成多模态基准、二值化 MNIST 和 LM1B 上,E-MoE 在少步生成方面优于因式分解基线。
查看 arXiv 页面
查看 PDF
添加到集合
社区
ArsenyIvanov
论文作者
论文提交者
1天前
掩码扩散语言模型并行生成多个标记,但其逆过程通常按位置进行因式分解,这限制了少步生成场景下的生成质量。我们引入了 E-MoE,它利用专家路由决策作为共享离散潜在变量,将混合专家主干网转化为因式分解分布的混合体。这使得协调的、非因式分解的生成成为可能。E-MoE 显著改善了语言建模中的少步生成,并在合成基准和二值化 MNIST 上恢复了因式分解基线未能捕捉到的多模态结构。
查看翻译
回复
librarian-bot
约21小时前
这是来自图书管理员机器人的自动消息。我发现了以下与本文相似的论文。
以下论文由 Semantic Scholar API 推荐
用于连续扩散语言模型的分布匹配蒸馏(2026)
LLaDA MoE v2:扩展混合专家扩散语言模型(2026)
离散扩散的单纯形松弛(2026)
SAGE:混合专家扩散模型中无分类器引导的子空间对齐(2026)
基于表示的掩码扩散模型(2026)
Alpha 扩散语言模型:因式分解本身并非问题(2026)
将线性注意力 retrofitting 到扩散语言模型中(2026)
如果您觉得此评论有帮助,请点赞!
如果您希望获取 Hugging Face 上任何论文的推荐,请查看此 Space
您可以通过在评论中标记它来直接向 Librarian Bot 请求论文推荐:@librarian-bot recommend
查看翻译
回复
编辑
预览
通过拖拽、粘贴或点击此处上传图片、音频和视频。
评论
· 注册或登录以发表评论
赞成票
57
+45
在您的智能体中获取此论文:
hf papers read 2609.37533
没有最新的 CLI?
引用此论文的模型
0
无链接此论文的模型
在模型的 README.md 中引用 arxiv.org/abs/2609.37533 以从本页面链接它。
引用此论文的数据集
0
无链接此论文的数据集
在数据集的 README.md 中引用 arxiv.org/abs/2609.37533 以从本页面链接它。
引用此论文的 Spaces
0
无链接此论文的 Space
在 Space 的 README.md 中引用 arxiv.org/abs/2609.37533 以从本页面链接它。
包含此论文的集合
0
无包含此论文的集合
将此论文添加到集合中以从本页面链接它。
Papers
arxiv:2609.37533
Copy markdown
E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models
Published on Sep 29
·
Submitted by
Arseny Ivanov
on Oct 2
Authors:
Arseny Ivanov
,
Alexander Kolesov
,
Alexander Korotin
,
Ivan Oseledets
,
Mikhail Goncharov
Abstract
Masked diffusion models (MDMs) generate sequences by progressively unmasking several tokens per denoising step, but their reverse process is typically factorized over positions, limiting sample quality in the few-step regime where diffusion's speed advantage over autoregressive decoding matters most. A recent line of work introduces a continuous Gaussian latent, trained as a variational autoencoder, to capture correlations across positions, but such approaches are prone to posterior collapse, where the latent is silently ignored. We propose Enhanced Mixture-of-Experts (E-MoE), which builds the reverse process as a mixture of factorized distributions over a discrete shared latent given by the expert-routing decisions of a Mixture-of-Experts (MoE) backbone, without increasing active parameters over the factorized baseline. Across synthetic multi-modal benchmarks, binarized MNIST, and LM1B, E-MoE improves few-step generation over factorized baselines.
View arXiv page
View PDF
Add to collection
Community
ArsenyIvanov
Paper author
Paper submitter
1 day ago
Masked diffusion language models generate multiple tokens in parallel, but their reverse process is typically factorized across positions, which limits generation quality in the few-step regime. We introduce E-MoE, which turns a Mixture-of-Experts backbone into a mixture of factorized distributions using expert-routing decisions as a shared discrete latent variable. This enables coordinated, non-factorized generation. E-MoE substantially improves few-step generation on language modeling and recovers multimodal structure that factorized baselines fail to capture on synthetic benchmarks and binarized MNIST.
See translation
Reply
librarian-bot
about 21 hours ago
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
Distribution Matching Distillation for Continuous Diffusion Language Models (2026)
LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models (2026)
Simplex Relaxation for Discrete Diffusion (2026)
SAGE: Subspace Alignment for Classifier-Free Guidance in Mixture-of-Experts Diffusion Models (2026)
Representation-based Masked Diffusion Model (2026)
Alpha Diffusion Language Models: Factorization Alone Is Not the Problem (2026)
Retrofitting Linear Attention into Diffusion Language Models (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
See translation
Reply
Edit
Preview
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Comment
· Sign up or log in to comment
Upvote
57
+45
Get this paper in your agent:
hf papers read 2609.37533
Don't have the latest CLI?
Models citing this paper
0
No model linking this paper
Cite arxiv.org/abs/2609.37533 in a model README.md to link it from this page.
Datasets citing this paper
0
No dataset linking this paper
Cite arxiv.org/abs/2609.37533 in a dataset README.md to link it from this page.
Spaces citing this paper
0
No Space linking this paper
Cite arxiv.org/abs/2609.37533 in a Space README.md to link it from this page.
Collections including this paper
0
No Collection including this paper
Add this paper to a collection to link it from this page.
| 刊期 | 得分 | 排名 | 结果 |
|---|---|---|---|
| 2026-10-04 | 10.22 | 14 | 入选 |
| 2026-10-03 | 10.95 | 13 | 未入选 |