# Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Digest AI · Research · published 2026-10-01T15:01:43Z · updated 2026-10-01T15:16:01Z

Canonical: https://digestai.news/story/introducing-olmo-core-3-open-scalable-training-infrastructure-for-larg

## Summary

The system aims to scale MoE models into the trillion-parameter range while maintaining computational efficiency. MoE models split learned components across experts, activating only a subset per input to reduce costs, but coordinating inputs across clusters adds overhead that can negate efficiency gains as models grow.

Olmo-core 3 addresses this by redesigning the training stack to minimize communication costs. In benchmarks, it increased the expert pool from **8 to 128** while keeping active parameters per token near **3.2B**, growing total capacity from **4.6B to 47B** parameters. Training throughput dropped by less than **5%** compared to earlier versions. The framework also supports **one trillion total parameters** in tests, with optimizations like expert parallelism, pipeline parallelism, and distributed optimizers to distribute workloads across GPUs. A preliminary test on **eight NVIDIA B300 GPUs** showed a **47-billion-parameter MoE** processed **52,000 tokens per second per GPU**, a **2.7×** improvement over the previous implementation.

The release includes techniques like **rowwise expert parallelism**, **GPU-resident routing**, and **grouped GEMM** to reduce data movement and computation overhead. It also supports **MXFP8**, a lower-precision format, which increased training throughput by **21%** in controlled tests while reducing peak memory usage from **103 GiB to 95 GiB**. The framework is fully open-source, allowing researchers to adapt it for different hardware and experiments.

## Key points

- Olmo-core 3 scales MoE training to **47B parameters** while cutting throughput loss to under **5%** compared to earlier versions
- Benchmark on **eight NVIDIA B300 GPUs** achieved **52,000 tokens/sec/GPU** for a **47B-parameter MoE**, a **2.7×** speedup over prior design
- Framework supports **MXFP8 precision**, boosting throughput by **21%** and reducing memory use from **103 GiB to 95 GiB** in tests

## Why it matters

Olmo-core 3 lowers the barrier for training massive MoE models by optimizing compute efficiency, enabling labs to explore larger architectures without proportional cost increases. Its open-source design accelerates research in sparse models, a key area for balancing performance and scalability.

## Sources

1. [Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs](https://huggingface.co/blog/allenai/olmocore3) (Hugging Face, 2026-10-01, primary source)
2. [Ai2 Releases Olmo-Core 3, Open Training Stack for Trillion-Parameter MoEs](https://unite.ai/ai2-releases-olmo-core-3-open-training-stack-for-trillion-parameter-moes) (Unite.AI, 2026-10-01)

## Cite

Digest AI, "Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs", 1 October 2026, https://digestai.news/story/introducing-olmo-core-3-open-scalable-training-infrastructure-for-larg

---

Written by Digest AI's editorial model from the linked sources; the sources are the record. Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse
JSON: https://digestai.news/story/introducing-olmo-core-3-open-scalable-training-infrastructure-for-larg.json
