Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs
The system aims to scale MoE models into the trillion-parameter range while maintaining computational efficiency. MoE models split learned components across experts, activating only a subset per input to reduce costs, but coordinating inputs across clusters adds overhead that can negate efficiency gains as models grow.
Key points
- Olmo-core 3 scales MoE training to **47B parameters** while cutting throughput loss to under **5%** compared to earlier versions
- Benchmark on **eight NVIDIA B300 GPUs** achieved **52,000 tokens/sec/GPU** for a **47B-parameter MoE**, a **2.7×** speedup over prior design
- Framework supports **MXFP8 precision**, boosting throughput by **21%** and reducing memory use from **103 GiB to 95 GiB** in tests
Olmo-core 3 addresses this by redesigning the training stack to minimize communication costs. In benchmarks, it increased the expert pool from 8 to 128 while keeping active parameters per token near 3.2B, growing total capacity from 4.6B to 47B parameters. Training throughput dropped by less than 5% compared to earlier versions. The framework also supports one trillion total parameters in tests, with optimizations like expert parallelism, pipeline parallelism, and distributed optimizers to distribute workloads across GPUs. A preliminary test on eight NVIDIA B300 GPUs showed a 47-billion-parameter MoE processed 52,000 tokens per second per GPU, a 2.7× improvement over the previous implementation.
The release includes techniques like rowwise expert parallelism, GPU-resident routing, and grouped GEMM to reduce data movement and computation overhead. It also supports MXFP8, a lower-precision format, which increased training throughput by 21% in controlled tests while reducing peak memory usage from 103 GiB to 95 GiB. The framework is fully open-source, allowing researchers to adapt it for different hardware and experiments.
Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs
Hugging Face · 1 October 2026
Loading the full article…
This text was published by Hugging Face. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- DeepEvidence agent explores biomedical evidence beyond retrieval · 1 src
- MIT researchers say ChatGPT can fuel delusional spiraling · 1 src
- NVIDIA releases Kumo Tabular model for tabular prediction with open weights · 2 src
- Researchers release ArgGYM benchmark for testing defeasible reasoning in AI models · 1 src
- Researchers release SimTrace for generating synthetic user behavior data · 1 src
Comments
via GitHub Discussions