DigestAI news desk

Cut through the AI noise.

Research6 min read

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

The system aims to scale MoE models into the trillion-parameter range while maintaining computational efficiency. MoE models split learned components across experts, activating only a subset per input to reduce costs, but coordinating inputs across clusters adds overhead that can negate efficiency gains as models grow.

1 source primary source

Key points

  • Olmo-core 3 scales MoE training to **47B parameters** while cutting throughput loss to under **5%** compared to earlier versions
  • Benchmark on **eight NVIDIA B300 GPUs** achieved **52,000 tokens/sec/GPU** for a **47B-parameter MoE**, a **2.7×** speedup over prior design
  • Framework supports **MXFP8 precision**, boosting throughput by **21%** and reducing memory use from **103 GiB to 95 GiB** in tests

Olmo-core 3 addresses this by redesigning the training stack to minimize communication costs. In benchmarks, it increased the expert pool from 8 to 128 while keeping active parameters per token near 3.2B, growing total capacity from 4.6B to 47B parameters. Training throughput dropped by less than 5% compared to earlier versions. The framework also supports one trillion total parameters in tests, with optimizations like expert parallelism, pipeline parallelism, and distributed optimizers to distribute workloads across GPUs. A preliminary test on eight NVIDIA B300 GPUs showed a 47-billion-parameter MoE processed 52,000 tokens per second per GPU, a 2.7× improvement over the previous implementation.

The release includes techniques like rowwise expert parallelism, GPU-resident routing, and grouped GEMM to reduce data movement and computation overhead. It also supports MXFP8, a lower-precision format, which increased training throughput by 21% in controlled tests while reducing peak memory usage from 103 GiB to 95 GiB. The framework is fully open-source, allowing researchers to adapt it for different hardware and experiments.

Full story from Hugging Face primary sourceOpen source ↗

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Hugging Face · 1 October 2026

Loading the full article…

This text was published by Hugging Face. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories