📄
科技报道 | 💻
代码 | 🧩
交互式演示
今天,我们发布 Olmo-core 3,这是我们对大型语言模型开发框架的重大升级,其特色是重新设计的开放混合专家(MoE)训练系统。
Olmo-core 3 旨在将 MoE 训练扩展至万亿参数规模,同时保持计算效率。它是下一代 Olmo 背后的核心系统之一,也是我们持续致力于公开每个新模型背后工具和训练基础设施承诺的一部分。
训练大型 AI 模型需要大量的计算资源,这不仅推高了成本,还增加了能源消耗,使得许多学术研究人员和小型实验室难以触及先进的模型开发。MoE 模型提供了一种更高效的途径——它们可以包含更多的学习组件(即参数),而无需让每个输入都使用所有参数。然而,整个模型仍然必须存储在 GPU 内存中并在训练期间进行更新,而在集群中将输入引导至正确的专家(MoE 内的专用组件)也会产生自身的通信和协调成本。随着 MoE 规模的扩大,这些成本可能会侵蚀仅对每个输入使用部分模型所带来的大部分计算优势。
Olmo-core 3 旨在填补这一差距。在一个基准测试中,我们将专家池从 8 个增加到 128 个,同时仍仅为每个 token(语言模型处理的小文本单元)选择四个专家,使每个 token 的活跃参数数量大致保持在约 32 亿。总参数容量从 46 亿增长到 470 亿,而训练吞吐量下降不到 5%。
同一基础设施在基准测试中已支持超过一万亿的总参数。
围绕 MoE 的实际工作原理构建训练栈
Olmo-core 随着每一代 Olmo 的演进而发展。
我们在稀疏模型方面的工作始于 OlmoE,它使用带有 64 个路由专家的 MoE 架构。相比之下,Olmo 3 使用了密集架构,这意味着几乎所有模型参数都对每个 token 处于活跃状态,其训练栈也是围绕该设计构建的。Olmo-core 3 通过一个专为更大规模 MoE 模型设计的训练系统扩展了该框架。
我们在 Olmo-core 中早期的 MoE 实现使用完全分片数据并行(FSDP),配置为收集并重新分片每个小批训练数据的模型权重。Olmo-core 3 转向基于分布式数据并行(DDP)的系统。它使专家驻留在 GPU 上并将相关数据路由到它们,从而避免了重复的权重收集。
NVIDIA 的 Megatron-Core 是训练大型 MoE 的既定选择。Olmo-core 3 为 Olmo 背后的框架带来了集成的 MoE 训练栈,其重新设计提高了吞吐量,优于我们早期基于 FSDP 的实现。在八块 NVIDIA B300 GPU 上的初步测试中,使用新栈的 470 亿参数 MoE 每块 GPU 每秒处理 52,000 个 token,而使用我们早期实现时仅为 19,400——吞吐量约为原来的 2.7 倍。
扩展和优化 MoE 训练
Olmo-core 3 结合了多种技术,用于在 GPU 集群上分布大型 MoE,并通过优化使路由和计算更加高效。
三种技术决定了模型及其训练状态如何在硬件上拆分:
专家并行(Expert parallelism)将专家分布在 GPU 上,因此每个 GPU 仅存储完整专家池的一部分。
流水线并行(Pipeline parallelism)将模型的层(转换输入的连续阶段)分布在 GPU 组之间,减少了每个 GPU 需要保留在内存中的模型部分。
分布式优化器(Distributed optimizer)将优化器状态(用于计算和应用训练期间更新所需的额外数据)分布在 GPU 上,而不是在每个 GPU 上存储完整副本。
这些技术共同作用,使得 MoE 能够扩展,而无需每个 GPU 都将整个模型及其训练状态保留在内存中。
Olmo-core 3 还降低了将数据路由到正确专家并运行其计算的成本。行级专家并行性将路由后的数据直接放入专家输入缓冲区,从而最大限度地减少重新排列数据所需的额外工作。GPU 驻留式路由将路由元数据保留在 GPU 上,因此 CPU 可以在无需等待信息复制回来的情况下排队处理工作。而分组 GEMM(通用矩阵乘法)则将许多小型专家计算组合在一起,使 GPU 能够更高效地执行它们。
最后,Olmo-core 3 支持 MXFP8,这是一种较低精度的数字格式,用更少的位数来表示某些数值。只要节省下来的成本超过在数字格式之间转换的成本,这就能减少计算量以及 GPU 之间移动的数据量。
我们在四块 NVIDIA B300 GPU 上进行的受控基准测试中测量了 MXFP8 对端到端训练吞吐量的影响,工作负载在专家之间均匀分布。在系统中最能受益的部分启用 MXFP8 后,与作为基线的较高精度格式 BF16 相比,训练吞吐量提高了约 21%,同时峰值活跃内存从 103 GiB 降至 95 GiB。大部分增益来自前馈计算以及在专家之间移动数据,而不仅仅是注意力机制。
这些技术和优化必须协同工作。加快训练的某一部分可能会在其他地方产生成本;更快的计算可能需要更多的数据移动,而如果转换数据耗时过长,减少位数可能无济于事。Olmo-core 3 围绕整个训练过程中的这些权衡而构建,使我们——以及使用开源堆栈的研究人员——能够控制各个部分如何组合在一起。
探索我们的交互式演示,了解数据、专家和流水线并行性如何协同工作以扩展 MoE 训练规模——从单块 GPU 到多块 GPU。
扩展至万亿参数范围
我们已在 NVIDIA B300 GPU 上的各种配置中对 Olmo-core 3 进行了基准测试,其中包括一个拥有 1.2 万亿参数的模型,在 512 块 GPU 上运行时,每个 token 有 583.6 亿个参数处于激活状态。其观察到的最高吞吐量为每块 GPU 858 TFLOP/s——这是衡量每块 GPU 每秒有用模型计算量的指标。这些测试使用随机路由来测量系统性能,而非已训练模型的质量。
我们还尝试了 DeepEP v2,这是一种处理 GPU 间专家通信的替代方法,达到了拥有 2.38 万亿总参数的配置。这是一项短期容量测试,而非完整的训练运行,因此它展示了 Olmo-core 3 所能达到的规模,而非持续的训练性能。
在这些规模下,系统性能只是图景的一部分。我们的技术报告还记录了指导我们如何训练 MoE 并衡量其性能的实验。例如:
旨在鼓励平衡路由的评分指标,在实际工作负载变得不那么平衡时甚至可能改善。我们将此称为失败 token 杰利蝾螈(gerrymandering)现象。
由于专家处理的 token 数量减少而降低其学习率——即训练更新的幅度——在我们测试的模型系列中并未带来结果提升。
当处理值发生变化时,即使矩阵维度相同,GPU 计算所花费的时间也不同。因此,性能比较需要匹配的输入值以及匹配的形状。
在单独的 GPU 流上重叠通信和计算并不总是能加快训练速度。在某些测试中,它反而减慢了端到端执行速度——这提醒我们,更多的重叠并不一定意味着更高的吞吐量。
报告解释了这些发现,以及我们测试但未采用的方法。
为下一代 Olmo 打造,向所有人开放
Olmo-core 3 是我们正在构建的下一代的基石。我们的下一代 Olmo 将采用 MoE 架构,我们的目标是使其成为迄今为止最强大的 Olmo,在最大的数据集和最长上下文窗口上进行训练。
新的技术栈使我们能够在之前的混合专家(MoE)工作基础上实现更广泛的扩展,同时随着模型和硬件的演进,为我们提供了更大的灵活性来调整训练策略。而且它是完全开源的——研究人员和开发者可以使用 Olmo-core 3 训练自己的 MoE 模型,将其适配到不同的硬件上,并对路由、并行化以及其他系统组件进行实验。
这正是我们看待开源模型开发的方式——当支撑模型的底层基础设施和训练决策也处于开放状态时,模型权重才更具价值。
若想深入了解系统设计、实验结果、消融研究以及我们在过程中测试的各种方法,请阅读我们的技术报告,并在 GitHub 上探索 Olmo-core 3。
📄
Tech Report | 💻
Code | 🧩
Interactive demo
Today we’re releasing Olmo-core 3 , a significant upgrade to our framework for developing large language models featuring a redesigned open mixture-of-experts (MoE) training system.
Olmo-core 3 is designed to scale MoE training into the trillion-parameter range while preserving computational efficiency. It’s one of the core systems behind the next generation of Olmo, and part of our ongoing commitment to open up the tools and training infrastructure behind each new model.
Training large AI models takes a lot of compute, driving up costs and energy use and putting advanced model development out of reach for many academic researchers and smaller labs. MoE models offer a more efficient approach—they can contain many more learned components, or parameters, without requiring every input to use all of them. But the full model still has to be stored across GPU memory and updated during training, and directing inputs to the right experts – the specialized components within an MoE – across a cluster creates its own communication and coordination costs. As MoEs grow, those costs can erode much of the computational advantage of using only part of the model for each input.
Olmo-core 3 is built to close that gap. In one benchmark, we increased the expert pool from 8 to 128 while still selecting only four experts per token – the small units of text a language model processes – keeping the number of active parameters per token roughly fixed at about 3.2B. Total parameter capacity grew from 4.6B to 47B, while training throughput fell by less than 5%.
The same infrastructure has been benchmarked at over one trillion total parameters.
Building a training stack around how MoEs actually work
Olmo-core has evolved with each generation of Olmo.
Our work on sparse models goes back to OlmoE , which used an MoE architecture with 64 routed experts. Olmo 3 , by contrast, used a dense architecture, meaning nearly all of the model was active for every token and its training stack was built around that design. Olmo-core 3 extends the framework with a training system designed for much larger MoE models.
Our earlier MoE implementation in Olmo-core used fully sharded data parallelism (FSDP), configured to gather and reshard model weights for each small batch of training data. Olmo-core 3 switches to a system based on distributed data parallelism (DDP) . It keeps experts resident on GPUs and routes the relevant data to them, avoiding that repeated weight gathering.
NVIDIA’s Megatron-Core is an established option for training large MoEs. Olmo-core 3 brings an integrated MoE training stack to the framework behind Olmo, with a redesign that improves throughput over our earlier FSDP-based implementation. In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using our earlier implementation—about 2.7× the throughput.
Scaling and optimizing MoE training
Olmo-core 3 combines several techniques for distributing large MoEs across GPU clusters with optimizations that make routing and computation more efficient.
Three techniques determine how the model and its training state are split across hardware:
Expert parallelism spreads the experts across GPUs, so each GPU stores only part of the full expert pool.
Pipeline parallelism splits the model’s layers – the successive stages that transform an input – across groups of GPUs, reducing how much of the model each GPU needs to keep in memory.
A distributed optimizer spreads the optimizer state – the additional data used to calculate and apply updates during training – across GPUs instead of storing a full copy on every GPU.
Together, these techniques allow an MoE to scale without requiring every GPU to keep the entire model and its training state in memory.
Olmo-core 3 also reduces the cost of routing data to the right experts and running their computations. Rowwise expert parallelism places routed data directly into expert input buffers, minimizing the extra work needed to rearrange it. GPU-resident routing keeps routing metadata on the GPUs, so the CPU can queue work without waiting for that information to be copied back. And grouped GEMM combines many small expert computations so GPUs can execute them more efficiently.
Finally, Olmo-core 3 supports MXFP8 , a lower-precision number format that represents some values with fewer bits. This can reduce computation and the amount of data moved between GPUs, as long as those savings outweigh the cost of converting between number formats.
We measured MXFP8’s effect on end-to-end training throughput in a controlled benchmark on four NVIDIA B300 GPUs, with work distributed uniformly across experts. With MXFP8 enabled across the parts of the system where it helped most, training throughput was about 21% higher than with BF16, the higher-precision format we used as our baseline, while peak active memory fell from 103 GiB to 95 GiB. Most of the gain came from feed-forward computation and moving data between experts rather than attention alone.
These techniques and optimizations have to work together. Speeding up one part of training can create costs elsewhere; faster computation may require more data movement, while moving fewer bits may not help if converting the data takes too long. Olmo-core 3 is built around those trade-offs across the full training process, giving us – and researchers using the open stack – control over how the pieces fit together.
Explore our interactive walkthrough to see how data, expert, and pipeline parallelism work together to scale MoE training—from a single GPU to many.
Scaling into the trillion-parameter range
We’ve benchmarked Olmo-core 3 across a range of configurations on NVIDIA B300 GPUs, including a 1.2-trillion-parameter model with 58.36 billion parameters active per token across 512 GPUs. Its highest observed throughput was 858 TFLOP/s/GPU—a measure of useful model computation per second on each GPU. These tests used random routing to measure system performance, rather than the quality of a trained model.
We’ve also experimented with DeepEP v2, an alternative way of handling communication between experts across GPUs, reaching a configuration with 2.38 trillion total parameters. This was a short-capacity test rather than a full training run, so it demonstrates the scale Olmo-core 3 can reach rather than sustained training performance.
At these scales, systems performance is only part of the picture. Our technical report also documents experiments that informed how we train MoEs and measure their performance. For example:
A score intended to encourage balanced routing could improve even as the actual workload became less balanced. We call this failure token gerrymandering .
Lowering experts’ learning rates – the size of their training updates – because they process fewer tokens did not improve results in the model family we tested.
GPU calculations took different amounts of time when the values being processed changed, even with the same matrix dimensions. Performance comparisons therefore need matching input values as well as matching shapes.
Overlapping communication and computation on separate GPU streams did not always make training faster. In some tests, it slowed end-to-end execution—a reminder that more overlap does not necessarily mean higher throughput.
The report explains these findings alongside the approaches we tested and chose not to adopt.
Built for the next generation of Olmo, open for everyone
Olmo-core 3 is the foundation for what we’re building next. Our next-generation Olmo will use an MoE architecture, and we’re aiming for it to be our most capable Olmo yet, trained on our largest dataset and with our longest context window.
The new stack lets us scale beyond our previous MoE work while giving us more flexibility to adapt training as models and hardware evolve. And it’s fully open—researchers and developers can use Olmo-core 3 to train their own MoEs, adapt it to different hardware, and experiment with routing, parallelism, and other parts of the system.
That’s part of how we think about open model development—model weights are more useful when the infrastructure and training decisions behind them are open too.
For a deeper look at the systems design, experiments, ablations, and approaches we tested along the way, read our technical report and explore Olmo-core 3 on GitHub .
首次收录 · 2026-10-02 · 10.6 分