{"version":1,"type":"story","url":"https://digestai.news/story/introducing-olmo-core-3-open-scalable-training-infrastructure-for-larg","json":"https://digestai.news/story/introducing-olmo-core-3-open-scalable-training-infrastructure-for-larg.json","markdown":"https://digestai.news/story/introducing-olmo-core-3-open-scalable-training-infrastructure-for-larg.md","slug":"introducing-olmo-core-3-open-scalable-training-infrastructure-for-larg","headline":"Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs","summary":"The system aims to scale MoE models into the trillion-parameter range while maintaining computational efficiency. MoE models split learned components across experts, activating only a subset per input to reduce costs, but coordinating inputs across clusters adds overhead that can negate efficiency gains as models grow.\n\nOlmo-core 3 addresses this by redesigning the training stack to minimize communication costs. In benchmarks, it increased the expert pool from **8 to 128** while keeping active parameters per token near **3.2B**, growing total capacity from **4.6B to 47B** parameters. Training throughput dropped by less than **5%** compared to earlier versions. The framework also supports **one trillion total parameters** in tests, with optimizations like expert parallelism, pipeline parallelism, and distributed optimizers to distribute workloads across GPUs. A preliminary test on **eight NVIDIA B300 GPUs** showed a **47-billion-parameter MoE** processed **52,000 tokens per second per GPU**, a **2.7×** improvement over the previous implementation.\n\nThe release includes techniques like **rowwise expert parallelism**, **GPU-resident routing**, and **grouped GEMM** to reduce data movement and computation overhead. It also supports **MXFP8**, a lower-precision format, which increased training throughput by **21%** in controlled tests while reducing peak memory usage from **103 GiB to 95 GiB**. The framework is fully open-source, allowing researchers to adapt it for different hardware and experiments.","keyPoints":["Olmo-core 3 scales MoE training to **47B parameters** while cutting throughput loss to under **5%** compared to earlier versions","Benchmark on **eight NVIDIA B300 GPUs** achieved **52,000 tokens/sec/GPU** for a **47B-parameter MoE**, a **2.7×** speedup over prior design","Framework supports **MXFP8 precision**, boosting throughput by **21%** and reducing memory use from **103 GiB to 95 GiB** in tests"],"whyItMatters":"Olmo-core 3 lowers the barrier for training massive MoE models by optimizing compute efficiency, enabling labs to explore larger architectures without proportional cost increases. Its open-source design accelerates research in sparse models, a key area for balancing performance and scalability.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":[],"models":[],"people":[]},"firstPublishedAt":"2026-10-01T15:01:43Z","updatedAt":"2026-10-01T15:16:01Z","sourceCount":2,"hasPrimarySource":true,"sources":[{"outlet":"Hugging Face","title":"Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs","url":"https://huggingface.co/blog/allenai/olmocore3","publishedAt":"2026-10-01T15:01:43Z","type":"primary","primary":true,"lead":true},{"outlet":"Unite.AI","title":"Ai2 Releases Olmo-Core 3, Open Training Stack for Trillion-Parameter MoEs","url":"https://unite.ai/ai2-releases-olmo-core-3-open-training-stack-for-trillion-parameter-moes","publishedAt":"2026-10-01T15:16:01Z","type":"press","primary":false,"lead":false}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs\", 1 October 2026, https://digestai.news/story/introducing-olmo-core-3-open-scalable-training-infrastructure-for-larg","publisher":"Digest AI","title":"Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs","datePublished":"2026-10-01T15:01:43Z","url":"https://digestai.news/story/introducing-olmo-core-3-open-scalable-training-infrastructure-for-larg"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}