In this article
AllenAI releases Olmo-core 3, a new infrastructure for training large MoE models
AllenAI has launched Olmo-core 3, a training framework designed to scale mixture-of-experts (MoE) models into the trillion-parameter range.
The system aims to maintain computational efficiency as model sizes grow. It forms part of the infrastructure behind the next generation of the Olmo series and continues the project’s commitment to open-sourcing the tools used to build these models.
Training large AI models requires significant compute, which drives up costs and energy use. This often excludes academic researchers and smaller labs from advanced model development. MoE models offer a more efficient approach by allowing many learned components without requiring every input to engage all of them. However, the full model must still reside in GPU memory. Directing inputs to the right experts within an MoE cluster creates communication and coordination costs. As MoEs grow, these costs can erode the advantage of using only part of the model for each input.
Olmo-core 3 addresses this gap. In one benchmark, the expert pool increased from 8 to 128 while selecting only four experts per token. This kept the number of active parameters per token roughly fixed at about 3.2 billion. Total parameter capacity grew from 4.6 billion to 47 billion, while training throughput fell by less than 5%.
The same infrastructure has been benchmarked at over one trillion total parameters.
Designing a training stack for sparse models
Olmo-core has evolved with each generation of Olmo.
Work on sparse models began with OlmoE, which used an MoE architecture with 64 routed experts. Olmo 3, by contrast, used a dense architecture where nearly all of the model was active for every token. Its training stack was built around that design. Olmo-core 3 extends the framework with a system designed for much larger MoE models.
Earlier MoE implementations in Olmo-core used fully sharded data parallelism (FSDP), configured to gather and reshard model weights for each small batch of training data. Olmo-core 3 switches to a system based on distributed data parallelism (DDP). It keeps experts resident on GPUs and routes the relevant data to them, avoiding repeated weight gathering.
NVIDIA’s Megatron-Core is an established option for training large MoEs. Olmo-core 3 brings an integrated MoE training stack to the framework behind Olmo. A redesign improves throughput over the earlier FSDP-based implementation. In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack. The earlier implementation achieved 19,400 tokens per second per GPU, representing about 2.7 times the throughput.
Scaling and optimizing MoE training
Olmo-core 3 combines several techniques for distributing large MoEs across GPU clusters with optimizations that make routing and computation more efficient.
Three techniques determine how the model and its training state are split across hardware:
- Expert parallelism spreads the experts across GPUs, so each GPU stores only part of the full expert pool.
- Pipeline parallelism splits the model’s layers across groups of GPUs, reducing how much of the model each GPU needs to keep in memory.
- A distributed optimizer spreads the optimizer state across GPUs instead of storing a full copy on every GPU.
Together, these techniques allow an MoE to scale without requiring every GPU to keep the entire model and its training state in memory.
Olmo-core 3 also reduces the cost of routing data to the right experts and running their computations. Rowwise expert parallelism places routed data directly into expert input buffers, minimising the extra work needed to rearrange it. GPU-resident routing keeps routing metadata on the GPUs, so the CPU can queue work without waiting for that information to be copied back. Grouped GEMM combines many small expert computations so GPUs can execute them more efficiently.
Finally, Olmo-core 3 supports MXFP8, a lower-precision number format that represents some values with fewer bits. This can reduce computation and the amount of data moved between GPUs, as long as those savings outweigh the cost of converting between number formats.
We measured MXFP8’s effect on end-to-end training throughput in a controlled benchmark on four NVIDIA B300 GPUs, with work distributed uniformly across experts. With MXFP8 enabled across the parts of the system where it helped most, training throughput was about 21% higher than with BF16, the higher-precision format used as a baseline. Peak active memory fell from 103 GiB to 95 GiB. Most of the gain came from feed-forward computation and moving data between experts rather than attention alone.
These techniques and optimizations must work together. Speeding up one part of training can create costs elsewhere. Faster computation may require more data movement, while moving fewer bits may not help if converting the data takes too long. Olmo-core 3 is built around those trade-offs across the full training process, giving researchers using the open stack control over how the pieces fit together.
Scaling into the trillion-parameter range
We have benchmarked Olmo-core 3 across a range of configurations on NVIDIA B300 GPUs. This includes a 1.2-trillion-parameter model with 58.36 billion parameters active per token across 512 GPUs. Its highest observed throughput was 858 TFLOP/s/GPU, a measure of useful model computation per second on each GPU. These tests used random routing to measure system performance, rather than the quality of a trained model.
We have also experimented with DeepEP v2, an alternative way of handling communication between experts across GPUs. This reached a configuration with 2.38 trillion total parameters. This was a short-capacity test rather than a full training run, so it demonstrates the scale Olmo-core 3 can reach rather than sustained training performance.
At these scales, systems performance is only part of the picture. Our technical report also documents experiments that informed how we train MoEs and measure their performance. For example:
- A score intended to encourage balanced routing could improve even as the actual workload became less balanced. We call this failure token gerrymandering.
- Lowering experts’ learning rates because they process fewer tokens did not improve results in the model family we tested.
- GPU calculations took different amounts of time when the values being processed changed, even with the same matrix dimensions. Performance comparisons therefore need matching input values as well as matching shapes.
- Overlapping communication and computation on separate GPU streams did not always make training faster. In some tests, it slowed end-to-end execution. More overlap does not necessarily mean higher throughput.
The report explains these findings alongside the approaches we tested and chose not to adopt.
Built for the next generation of Olmo
Olmo-core 3 is the foundation for what we are building next. Our next-generation Olmo will use an MoE architecture and aims to be the most capable model yet, trained on the largest dataset with the longest context window.
The new stack lets us scale beyond previous MoE work while giving more flexibility to adapt training as models and hardware evolve. It is fully open. Researchers and developers can use Olmo-core 3 to train their own MoEs, adapt it to different hardware, and experiment with routing, parallelism, and other parts of the system.
This reflects the approach to open model development: model weights are more useful when the infrastructure and training decisions behind them are open too.
Read the technical report and explore Olmo-core 3 on GitHub for a deeper look at the systems design, experiments, and approaches tested.




