论文
arxiv:2609.21346
复制 Markdown
IntBMoE:将块级条件控制整合到专家组合中以实现全参与混合专家模型
发表于 9月18日
·
提交者
徐龙飞
于 9月21日
作者:
程然
,
徐龙飞
,
刘正
,
刘凯魁
,
褚祥祥
摘要
混合专家(MoE)模型能够扩展容量,但现有设计无法独立设置三个关键数量。对于单个令牌而言,“参与度”是指有多少个专家为其输出贡献知识,“执行量”是指实际计算了多少个专家(计算成本),而“实例化量”是指必须构建和存储多少个专家大小的参数集(内存成本)。稀疏路由保持了较低的执行量和实例化量,但降低了参与度:对于每个令牌,仅有少数专家做出贡献。密集输出混合恢复了全参与度,但其执行量随专家数量的增加而增长。参数合并将执行量保持在一个专家的水平,但其实例化量随路由决策的数量增加。我们提出了 IntBMoE,这是一种块条件控制的 MoE 模型,通过将密集专家组合与稀疏块执行相结合,解耦了上述三个指标。其“块”来自一个小型的学习代码本,每个入口对应一个。在每一层内部,一个轻量级的超网络将该层池中的所有专家基线合并为一个组合专家。参与度是全量的,因为每个组合专家都利用了整个池。执行量保持稀疏,因为路由器将每个令牌仅发送给少数几个块。实例化量是有界的,因为代码本(而非输入)决定了存在的块的数量。双路径残差门控(DPRG)进一步通过乘法门控耦合了两个独立组成的路径。在图像分类上的实验表明,其表现持续优于代表性的稀疏和密集 MoE 基线模型。在语言建模和序列推荐上的额外实验验证了其在视觉领域之外的泛化能力。IntBMoE 已完全部署于高德地图的生成式推荐系统中,在60毫秒的延迟预算下服务于数亿用户,在线 A/B 测试中相对 UVCTR 提升了2.4%。我们的代码可在 https://github.com/AMAP-ML/DreamX-Rec/ 获取。
查看 arXiv 页面
查看 PDF
添加到收藏
社区
Xufew
论文作者
论文提交者
大约11小时前
IntBMoE 通过结合共享专家池中的参数来学习可重用的计算块,在实现全专家参与度的同时,将每个令牌仅路由到少数几个块。这些块可以预先计算并缓存,当块配置固定时,使得推理计算独立于专家数量。在计算机视觉、自然语言处理和推荐领域的实验表明,其具有强大的预测性能和高效的推理能力,突显了其在不同领域中的有效性。
查看翻译
回复
researchstudio-bot
大约8小时前
这是来自 ResearchStudio 团队的自动消息。
我们为此论文创建了一个交互式的 ResearchStudio Reel。它包括可视化海报、视频和博客,所有文件均可下载且格式可编辑。
打开 ResearchStudio Reel →
从 Hugging Face 下载所有文件
如果您觉得这个 Reel 有帮助,请给此评论点赞!
想要探索或为更多论文创建 Reel?请访问 ResearchStudio 演示页面。
查看翻译
回复
编辑
预览
通过拖拽、粘贴或点击此处上传图片、音频和视频。
评论
· 注册或登录以发表评论
赞同
90
+78
在您的代理中获取此论文:
hf papers read 2609.21346
没有最新的 CLI?
引用此论文的模型
0
无链接此论文的模型
在模型的 README.md 中引用 arxiv.org/abs/2609.21346 以从此页面链接它。
引用此论文的数据集
0
无链接此论文的数据集
在数据集的 README.md 中引用 arxiv.org/abs/2609.21346 以从此页面链接它。
引用此论文的 Spaces
0
无链接此论文的 Space
在 Space 的 README.md 中引用 arxiv.org/abs/2609.21346 以从此页面链接它。
在 Space 的 README.md 中引用 arxiv.org/abs/2609.21346,以便从本页面链接到该论文。
包含此论文的集合
0
没有包含此论文的集合
将这篇论文添加到某个集合中,以便从本页面链接到它。
Papers
arxiv:2609.21346
Copy markdown
IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts
Published on Sep 18
·
Submitted by
LONGFEI XU
on Sep 21
Authors:
Ran Cheng
,
Longfei Xu
,
Zheng Liu
,
Kaikui Liu
,
Xiangxiang Chu
Abstract
Mixture-of-Experts (MoE) scales capacity, but existing designs cannot set three quantities independently. For a single token, participation is how many experts contribute knowledge to its output, execution is how many are actually computed (compute cost), and materialization is how many expert-sized parameter sets must be built and stored (memory cost). Sparse routing keeps execution and materialization low, but shrinks participation: for each token, only a few experts contribute. Dense output-mixing restores full participation, but its execution grows with the number of experts. Parameter-merging keeps execution at one expert, but its materialization grows with the number of routing decisions. We propose IntBMoE, a block-conditioned MoE that decouples all three by pairing dense expert composition with sparse block execution. Its blocks come from a small learned codebook, one per entry. At each internal layer, a lightweight hypernetwork merges all expert bases in that layer's pool into one composed expert. Participation is full, because every composed expert draws on the entire pool. Execution stays sparse, because a router sends each token to only a few blocks. Materialization is bounded, because the codebook, not the input, fixes how many blocks exist. Dual-Path Residual Gating (DPRG) further couples two independently composed paths through multiplicative gating. Experiments on image classification show consistent gains over representative sparse and dense MoE baselines. Additional experiments on language modeling and sequential recommendation validate its generalization beyond vision. IntBMoE is fully deployed in AMap's generative recommendation system, serving hundreds of millions of users under a 60ms latency budget, with a 2.4% relative UVCTR gain in online A/B testing. Our code is available at https://github.com/AMAP-ML/DreamX-Rec/.
View arXiv page
View PDF
Add to collection
Community
Xufew
Paper author
Paper submitter
about 11 hours ago
IntBMoE learns reusable computational blocks by combining parameters from a shared expert pool, enabling full expert participation while routing each token to only a few blocks. These blocks can be precomputed and cached, making inference computation independent of the expert count when the block configuration is fixed. Experiments across computer vision, natural language processing, and recommendation demonstrate strong predictive performance and efficient inference, highlighting its effectiveness across diverse domains.
See translation
Reply
researchstudio-bot
about 8 hours ago
This is an automated message from the ResearchStudio team.
We created an interactive ResearchStudio Reel for this paper. It includes a visual poster, a video, and a blog, all available for download in editable formats.
Open the ResearchStudio Reel →
Download all files from Hugging Face
Please give this comment a thumbs up if you find the Reel helpful!
Want to explore or create Reels for more papers? Visit the ResearchStudio demo.
See translation
Reply
Edit
Preview
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Comment
· Sign up or log in to comment
Upvote
90
+78
Get this paper in your agent:
hf papers read 2609.21346
Don't have the latest CLI?
Models citing this paper
0
No model linking this paper
Cite arxiv.org/abs/2609.21346 in a model README.md to link it from this page.
Datasets citing this paper
0
No dataset linking this paper
Cite arxiv.org/abs/2609.21346 in a dataset README.md to link it from this page.
Spaces citing this paper
0
No Space linking this paper
Cite arxiv.org/abs/2609.21346 in a Space README.md to link it from this page.
Collections including this paper
0
No Collection including this paper
Add this paper to a collection to link it from this page.
首次收录 · 2026-09-22 · 10.95 分