Xiaomi releases MiMo-V2.6-Pro-RL with 1M-token context and 1.02T parameters
Xiaomi introduced MiMo-V2.6-Pro-RL, the flagship checkpoint of the MiMo-V2.6 series, built to scale reinforcement learning toward self‑improvement. The model is native omnimodal, handling text, image, video, and audio, and supports a 1M‑token context window for long‑horizon tasks.
Key points
- MiMo-V2.6-Pro-RL supports text, image, video, audio with 1 M token context length.
- Model has 1.02 T total parameters, 42 B activated, using Sparse MoE architecture.
- RL training uses 1,568 prompts × 16 rollouts per step and groupwise agentic grading for self‑improvement.
The architecture uses a Sparse Mixture of Experts with 1.02 T total parameters and 42 B activated parameters. It includes a 681M‑param MiMo ViT vision encoder, a 308M AudioTokenizer plus a 127M audio patch encoder, and a 5‑layer speculative decoder. Training employed Group Relative Policy Optimization on very large batches of 1,568 prompts × 16 rollouts per step, with billions of tokens per update, and a groupwise agentic grading loop that ranks rollouts to drive self‑improvement. Deployment instructions reference SGLang and vLLM Docker images.
Model pages: MiMo-V2.6-Pro-RL → · MiMo-V2.6-Distill-Qwen-9B → · MiMo-V2.6-Flash-RL →
XiaomiMiMo/MiMo-V2.6-Pro-RL · Hugging Face
huggingface.co · 21 September 2026
Loading the full article…
This text was published by huggingface.co. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
Coverage and discussion
3sources- Reddit discussionreddit.com
- Reddit discussionreddit.com
- Reddit discussionreddit.com
- XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9BPrimary source · huggingface.co ·
- XiaomiMiMo/MiMo-V2.6-Flash-RL · Hugging FacePrimary source · huggingface.co ·
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
More in Generative AI & Models
All →- xAI launches Grok 4.7 for coding and knowledge work at $2 input and $6 output tokens · 4 src
- Alibaba's Qwen-Image-2.1 claims to beat closed models in image generation · 2 src
- Anthropic cuts Claude Fable 5.1 cache read costs by 75% · 1 src
- OpenAI classifies GPT-6 Astra as Critical for cybersecurity · 1 src
- Claude cannot forecast crypto or stock prices, benchmark shows humans outperform · 1 src
Comments
via GitHub Discussions