Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs
Olmo-core 3 is a significant upgrade to the framework behind the Olmo model family, featuring a redesigned open MoE training system built to scale to trillion-parameter models while preserving computational efficiency. It is one of the core systems behind the next generation of Olmo, released alongside a tech report, code, and an interactive demo, as part of an effort to open up the training infrastructure behind each new model. The motivation is cost and access: training large AI models takes enormous compute, driving up costs and energy use and putting advanced model development out of reach for many academic researchers and smaller labs. MoEs help because they hold many more parameters without requiring every input to use all of them, but the full model must still be stored across GPU memory and updated during training, and routing inputs to the right experts across a cluster adds communication and coordination costs that can erode the efficiency advantage as MoEs grow.
Architecturally, Olmo-core's earlier MoE implementation used fully sharded data parallelism (FSDP), configured to gather and reshard model weights for each small batch of training data. Olmo-core 3 instead uses a distributed data parallelism (DDP) system that keeps experts resident on GPUs and routes the relevant data to them, avoiding repeated weight gathering. The work builds on prior sparse-model efforts — OlmoE used an MoE architecture with 64 routed experts, while Olmo 3 used a dense architecture whose training stack was built around nearly all parameters being active per token. NVIDIA's Megatron-Core is noted as an established option for training large MoEs; Olmo-core 3 brings an integrated MoE training stack to the Olmo framework with a redesign that improves throughput over the earlier FSDP-based implementation.
In one benchmark, the expert pool grew from 8 to 128 while still selecting only four experts per token, keeping active parameters per token roughly fixed at about 3.2B. Total parameter capacity rose from 4.6B to 47B while training throughput fell by less than 5%. The same infrastructure has been benchmarked at over one trillion total parameters. In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, versus 19,400 with the earlier implementation — about 2.7× the throughput.