# Xiaomi releases MiMo-V2.6-Pro-RL with 1M-token context and 1.02T parameters

Digest AI · Generative AI & Models · published 2026-09-21T20:06:48Z · updated 2026-09-21T20:51:40Z

Canonical: https://digestai.news/story/xiaomi-releases-mimo-v2-6-pro-rl-with-1m-token-context-and-1-02t-param

## Summary

Xiaomi introduced MiMo-V2.6-Pro-RL, the flagship checkpoint of the MiMo-V2.6 series, built to scale reinforcement learning toward self‑improvement. The model is native omnimodal, handling text, image, video, and audio, and supports a 1M‑token context window for long‑horizon tasks.

The architecture uses a Sparse Mixture of Experts with 1.02 T total parameters and 42 B activated parameters. It includes a 681M‑param MiMo ViT vision encoder, a 308M AudioTokenizer plus a 127M audio patch encoder, and a 5‑layer speculative decoder. Training employed Group Relative Policy Optimization on very large batches of 1,568 prompts × 16 rollouts per step, with billions of tokens per update, and a groupwise agentic grading loop that ranks rollouts to drive self‑improvement. Deployment instructions reference SGLang and vLLM Docker images.

## Key points

- MiMo-V2.6-Pro-RL supports text, image, video, audio with 1 M token context length.
- Model has 1.02 T total parameters, 42 B activated, using Sparse MoE architecture.
- RL training uses 1,568 prompts × 16 rollouts per step and groupwise agentic grading for self‑improvement.

## Why it matters

A trillion‑parameter, multimodal model that integrates large‑scale reinforcement learning and self‑improvement loops could accelerate development of autonomous agents and complex task solving.

## Sources

1. [XiaomiMiMo/MiMo-V2.6-Pro-RL · Hugging Face](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL) (huggingface.co, 2026-09-21, primary source)
2. [XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B) (huggingface.co, 2026-09-21, primary source)
3. [XiaomiMiMo/MiMo-V2.6-Flash-RL · Hugging Face](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL) (huggingface.co, 2026-09-21, primary source)

## Cite

Digest AI, "Xiaomi releases MiMo-V2.6-Pro-RL with 1M-token context and 1.02T parameters", 21 September 2026, https://digestai.news/story/xiaomi-releases-mimo-v2-6-pro-rl-with-1m-token-context-and-1-02t-param

---

Written by Digest AI's editorial model from the linked sources; the sources are the record. Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse
JSON: https://digestai.news/story/xiaomi-releases-mimo-v2-6-pro-rl-with-1m-token-context-and-1-02t-param.json
