Olmo-core 3 Released for Open MoE Training
The Allen Institute for AI has launched Olmo-core 3, introducing a redesigned open mixture-of-experts training infrastructure built to scale large language models into the trillion-parameter range while preserving computational efficiency.

Introduction to Olmo-core 3
The development team behind open artificial intelligence models has released Olmo-core 3, a significant upgrade to their framework for developing large language models featuring a redesigned open mixture-of-experts (MoE) training system. According to details shared on the Hugging Face Blog, the new infrastructure is designed to scale MoE training into the trillion-parameter range while preserving computational efficiency.
Training large AI models demands substantial compute resources, driving up both costs and energy use, which often puts advanced model development out of reach for academic researchers and smaller laboratories. While MoE models offer a more efficient approach by containing many more learned components or parameters without requiring every input to use all of them, storing the full model across GPU memory and directing inputs to the right specialized components creates communication and coordination costs that can erode computational advantages.

Architectural Evolution and Throughput Gains
Olmo-core 3 shifts away from earlier implementations that relied on fully sharded data parallelism (FSDP). Instead, the updated framework utilizes a system based on distributed data parallelism (DDP) that keeps experts resident on GPUs and routes relevant data directly to them, thereby avoiding repeated weight gathering operations.
In preliminary tests conducted on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU utilizing the new stack, compared to 19,400 tokens using the earlier implementation. This represents approximately 2.7 times the throughput. Further architectural insights, benchmark data, and technical specifications are detailed in the official Tech Report published by the research team.
Scaling Techniques and GPU Optimization
To effectively distribute large MoEs across GPU clusters, Olmo-core 3 combines several core techniques. Expert parallelism spreads the expert pool across multiple GPUs, pipeline parallelism splits model layers across groups of GPUs, and a distributed optimizer spreads the optimizer state rather than storing a full copy on every device. Developers interested in exploring the codebase can review the source repository via the project's Code link.
Additionally, the framework reduces the cost of routing data and running computations through rowwise expert parallelism, GPU-resident routing, and grouped GEMM execution. These combined optimizations allow massive MoE architectures to scale without requiring every single GPU to hold the entire model and its associated training state in memory simultaneously.

Precision Formats and Trillion-Parameter Benchmarks
Olmo-core 3 also introduces support for MXFP8, a lower-precision number format designed to reduce computation overhead and minimize data movement between GPUs. Controlled benchmarks on four NVIDIA B300 GPUs demonstrated that enabling MXFP8 across optimized system components yielded a training throughput roughly 21% higher than the baseline BF16 format, while peak active memory decreased from 103 GiB to 95 GiB.
The framework has been thoroughly benchmarked across various configurations on NVIDIA B300 hardware, including a massive 1.2-trillion-parameter model operating across 512 GPUs. Researchers can also engage with an Interactive demo to visualize how data, expert, and pipeline parallelism interact during large-scale MoE training configurations.
Sources
- Hugging Face BlogIntroducing Olmo-core 3: Open, scalable training infrastructure for large MoEs