3 Sources
[1]
'Better than DeepSeek': Xiaomi's MiMo-V2.6-Pro debuts as the top open weights model in the world alongside cheaper V2.6-Flash
In a surprising upset, Chinese electric car and consumer electronics manufacturer Xiaomi has released the latest version of its growing family of MiMo language models, and MiMo-V2.6-Pro has arrived as the top-performing open-weight model in the world on third-party benchmarking firm Artificial
[2]
Flash Variant of Xiaomi MiMo 2.6 Pro is Free on Open Router
The Xiaomi MiMo 2.6 Pro has emerged as a standout in the open source AI landscape, blending advanced reinforcement learning techniques with practical features designed for diverse applications. One of its most notable capabilities is the 1 million token context window, which enables the model to
[3]
Xiaomi releases MiMo-V2.6-Pro and MiMo-V2.6-Flash open-source AI models with scaled RL training
Xiaomi has released and open-sourced the MiMo-V2.6 series. The company says the release focuses on scaling reinforcement learning (RL) compute on verifiable and complex tasks through exploration and feedback. The series includes MiMo-V2.6-Pro and MiMo-V2.6-Flash, while Xiaomi is also rolling out
Share
Copy Link
Xiaomi's MiMo-V2.6-Pro debuts as the top-performing open weights model globally, scoring 46 on Artificial Analysis Intelligence Index. The MIT-licensed model features advanced reinforcement learning, 1 million token context window, and costs just $0.435 per million input tokens—dramatically undercutting proprietary competitors while delivering frontier-level performance.
Xiaomi has released MiMo-V2.6-Pro as the world's top-performing open weights model, achieving a score of 46 on Artificial Analysis Intelligence Index
1
. This positions the Chinese manufacturer's flagship ahead of proprietary models including xAI's Grok 4.6 at 44 and Google's Gemini 3.8 Flash at 41. The result ties MiMo-V2.6-Pro with the newly released Grok 4.7 and places it above DeepSeek V4.1 Flash at 39 and DeepSeek V4.1 Pro at 361
. The achievement marks a shift for Xiaomi, better known internationally for smartphones and electric vehicles than frontier AI development.
Source: Geeky Gadgets
The MiMo-V2.6-Pro is MIT-licensed and available for free download on Hugging Face, allowing developers to customize, fine-tune, and deploy it without paying Xiaomi
1
. Through Xiaomi's API, the model costs $0.435 per million uncached input tokens and $0.87 per million output tokens. Artificial Analysis measures it at $0.13 per Intelligence Index task with output speed at roughly 134 tokens per second1
. Alongside the flagship, Xiaomi launched MiMo-V2.6-Flash at $0.14 per million input tokens and $0.28 per million output tokens, making it the second cheapest major frontier model available globally1
. The Flash variant is available for free through Open Router for developers working with limited budgets2
.The breakthrough behind MiMo-V2.6 stems from scaled RL training across verifiable and complex tasks
3
. Xiaomi streamed the production run live, completing 30 RL steps across roughly 750,000 trajectories in under six days3
. Training costs reached approximately $0.85 million for Flash and $2.62 million for Pro3
. The average pass rate increased by 25% for Flash and 12% for Pro in relative terms. On DeepSWE v1.1, a held-out long-horizon software engineering benchmark, scores jumped from 48.8 to 65.68 for Flash and from 58.4 to 72.57 for Pro3
. The company scaled RL compute through larger batches using a fully asynchronous architecture with 1,568 samples per update, processing 3.5-3.7 billion tokens per step3
.Xiaomi's release represents another step toward building a complete open agent stack encompassing foundation models, coding agents, harnesses, and training infrastructure
1
. The company introduced MiMo Code in June as an open-source terminal coding agent with persistent cross-session memory and task checkpoints. Internal testing suggested MiMo Code's advantage over Claude Code became more pronounced after workflows stretched beyond 200 execution steps1
. Xiaomi also released HarnessX, a research framework treating prompts, memory systems, tools, and control logic as components that can be rewritten and optimized. The company reported an average 14.5% absolute performance gain across 15 model-benchmark combinations when the harness evolved dynamically1
. Xiaomi is open-sourcing the technical report, training environments, and RL code for reproduction and verification3
.Both models are natively omnimodal, supporting a 1 million token context window with text, image, audio, and video input
1
2
. The models excel in 3D spatial reasoning, generating detailed and accurate 3D environments for game development, virtual reality, and simulation projects2
. For robotics control, MiMo-V2.6-Pro enhances precision and adaptability in autonomous systems, supporting applications in manufacturing and logistics2
. The model can take an image, video, or text prompt, break it into tasks, and coordinate multiple agents to build 3D scenes, perform visual verification, and produce runnable interactive worlds3
. It can generate 3D objects and scenes in Blender from text descriptions or reference images for animation, 3D printing, and game development3
.Related Stories
Xiaomi is releasing MiMo-V2.6-Pro-UltraSpeed, which generates outputs up to 20 times faster than the standard Pro version
1
2
. This Pro Ultra Speed mode significantly reduces processing times for time-sensitive workflows. The architecture builds on the sparse mixture-of-experts approach established in MiMo-V2.5-Pro, a 1.02-trillion-parameter model with 42 billion parameters active during inference1
. These agentic models were trained specifically for long-horizon software engineering with harness awareness—the ability to manage memory and context while operating inside agent scaffolds over hundreds or thousands of tool calls1
. The Flash variant targets high-volume production workloads while retaining the same 1 million token context window and native multimodal capabilities1
.As training scaled, Xiaomi froze the router to suppress training drift and implemented defenses against reward hacking
3
. These measures include reward design, adversarial evaluation, anomaly detection, and cross-checking between verifiers to improve training stability and reward reliability3
. For agentic RL across different tasks, Xiaomi developed a unified trajectory representation and penalty mechanism, high-concurrency interaction across agent frameworks, a decoupled control and data plane, stabilized per-task sampling ratios, and optimized training and inference engines3
. The training suite covers coding, general agents, visual, and cyber tasks across several harnesses3
. Watch how Xiaomi's approach to grader compute provides more precise and diverse reward signals for long-horizon RL tasks, helping reduce paths and tokens per task while maintaining training effectiveness across the expanding model family.Summarized by
Navi
[1]
[2]
02 Jun 2026•Technology

04 Aug 2026•Technology

30 Apr 2025•Technology

1
Technology

2
Science and Research

3
Technology
