Announced 8 Oct 2026 · Sources checked
What did Xiaomi post on 8 October?
arXiv’s API lists 2610.11959 as published 8 October 2026 13:41 UTC. The abstract presents MiMo-V2.6 as an omni-modal family that scales RL compute along batch/throughput, environment diversity, and grader compute, with MoE-router freezing and anti–reward-hacking defenses.
The report’s open-source pitch includes MiMo-V2.6-Distill-Qwen-9B, task environments, and an RL framework. For license hygiene when grabbing Hub weights, keep open weights vs open source in view—and compare decision-model drops like Nace Drex v1.5 carefully; Drex is a different stack.
A public training-metrics site at https://mimo.xiaomi.com/rl/mimo-v26 loaded when we fetched it, describing live trainer logs for Pro and Flash RL runs. We did not audit log authenticity beyond HTTP retrieval.
The MiMo-V2.6 materials are a technical report plus whatever checkpoints Xiaomi actually uploaded. Separate those objects: a PDF can describe RL recipes without shipping every intermediate reward model. Cite the dated Hub revision you download, not only the blog’s marketing name.

How big are Pro and Flash, on Xiaomi’s numbers?
Inside the HTML report we opened, Xiaomi describes MiMo-V2.6-Pro as a 1.02T-parameter Mixture-of-Experts model with 42B active parameters, and MiMo-V2.6-Flash as 310B parameters with 15B active.
Architecture notes include hybrid Sliding Window Attention interleaved with Global Attention, a 128-token sliding window, multimodal encoders, and MTP blocks. Those are report specifications, not our measurements.
If you need a primer on MoE routing before reading the tables, see mixture of experts explained.
Pro versus Flash sizing is Xiaomi’s product split for quality versus latency/cost. When they quote agentic or coding benches, keep the harness, tool budget, and thinking-token policy in the same sentence as the score. A Flash win on latency with a Pro win on hard tasks is a portfolio result, not a contradiction.
What is actually open for other teams?
The Hugging Face model card for XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B describes a 9B agentic SFT checkpoint on Qwen3.5-9B for coding, general agents, visual coding, and cybersecurity, released as a starting point for open agentic RL research.
Hub API metadata we fetched listed license:mit, createdAt 2026-09-21, and lastModified 2026-09-22—so the small open checkpoint is older than the 8 October report even though the report points readers to it.
The card’s evaluation table compares Qwen3.5-9B with the Distill SFT checkpoint on SWE Verified, SWE Pro, AutomationBench, Terminal Bench 2.1, and internal MiMo mini sets. Figures such as SWE Verified 61.1 avg@3 are Xiaomi’s card numbers. We did not run them.

What RL ideas does the report emphasize?
Xiaomi highlights asynchronous training consuming 1,568 samples and about 2.7–3.7B tokens per step at context lengths up to 1M; mixed environments spanning code, general, visual, and cyber domains; and groupwise agentic grading with Groupwise Reward Synthesis and Advantage Redistribution.
Stability measures named in the abstract include freezing the MoE router and multi-layer defenses against reward hacking. Infrastructure notes mention unified trajectories, multi-framework rollout pools, and train/inference consistency for MoE routing.
Those mechanisms are described as Xiaomi’s production RL stack. Open-sourcing a framework is not the same as reproducing Pro-scale runs on academic budgets.
Xiaomi reports gains not only on verifiable coding tasks such as DeepSWE but also on less verifiable workflows such as web development and on internal visual and cyber benches. Domain breadth is part of the scaling claim; it is also why independent replication will take more than one unit test suite.
The Distill-Qwen-9B card’s quickstart uses a recent SGLang build with a mimo reasoning parser and an explicit enable_thinking chat-template flag. That is operational detail for anyone who does load the MIT checkpoint—we still did not.
Scaled agentic RL usually implies environments with tools, verifiers, and long rollouts. Note which environments Xiaomi says it used, whether rewards are unit tests versus model judges, and how they clipped or normalized advantages. Those choices transfer more cleanly than a single leaderboard number.
What should practitioners do with this drop?
If you want a small MIT-licensed agentic SFT base, Distill-Qwen-9B is the concrete artifact on the Hub, with an SGLang serve example on the card. Confirm template and reasoning-parser flags before eval.
If you care about the scaling recipe, read the 8 October report and the metrics site, and treat Pro/Flash benchmark curves as vendor research until third parties replicate them.
Do not conflate Distill SFT numbers with post-RL Pro/Flash claims. The card’s table is explicitly the released SFT checkpoint.
Watch whether Xiaomi keeps the metrics site updated as runs continue; a frozen curve screenshot is weaker evidence than a live log page, and neither replaces a third-party eval harness.
If weights are open enough for your license constraints, run your own agent suite before swapping production routers. If only the report is open, treat it as a methods note for your RL team and wait for a reproducible training config before budgeting a from-scratch rerun.
What did we not test?
We did not download safetensors, serve SGLang, run SWE-bench, or train with Xiaomi’s RL framework. Parameter counts, token throughput, and leaderboard rows above are Xiaomi’s.
Common questions
Are the giant Pro/Flash weights on Hugging Face?
The report and Hub search show multiple MiMo-V2.6-related repos; the card we inspected in detail is Distill-Qwen-9B (MIT). Treat large RL checkpoints’ access terms as separate from that 9B SFT card.
Why mention September if the report is 8 October?
Because the Distill-Qwen-9B Hub timestamps we fetched are 21–22 September. The technical report’s arXiv timestamp is 8 October. Keep those dates distinct.
Did AiLookout train MiMo?
No. We opened the arXiv paper, Hub card, and metrics page only.
What to remember
MiMo-V2.6’s 8 October report is Xiaomi’s account of scaled agentic RL—use Hub revisions for what is downloadable, and keep bench rows in Xiaomi’s harness column.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





