Xiaomi's MiMo-V2.6-Pro debuts as the top-performing open weights model globally, scoring 46 on Artificial Analysis Intelligence Index. The MIT-licensed model features advanced reinforcement learning, 1 million token context window, and costs just $0.435 per million input tokens—dramatically undercutting proprietary competitors while delivering frontier-level performance.

Xiaomi MiMo-V2.6-Pro Claims Top Spot Among Open Weights Models

Xiaomi has released MiMo-V2.6-Pro as the world's top-performing open weights model, achieving a score of 46 on Artificial Analysis Intelligence Index

1

. This positions the Chinese manufacturer's flagship ahead of proprietary models including xAI's Grok 4.6 at 44 and Google's Gemini 3.8 Flash at 41. The result ties MiMo-V2.6-Pro with the newly released Grok 4.7 and places it above DeepSeek V4.1 Flash at 39 and DeepSeek V4.1 Pro at 36

1

. The achievement marks a shift for Xiaomi, better known internationally for smartphones and electric vehicles than frontier AI development.

Source: Geeky Gadgets

Source: Geeky Gadgets

Open-Source AI Models with Aggressive API Pricing

The MiMo-V2.6-Pro is MIT-licensed and available for free download on Hugging Face, allowing developers to customize, fine-tune, and deploy it without paying Xiaomi

1

. Through Xiaomi's API, the model costs $0.435 per million uncached input tokens and $0.87 per million output tokens. Artificial Analysis measures it at $0.13 per Intelligence Index task with output speed at roughly 134 tokens per second

1

. Alongside the flagship, Xiaomi launched MiMo-V2.6-Flash at $0.14 per million input tokens and $0.28 per million output tokens, making it the second cheapest major frontier model available globally

1

. The Flash variant is available for free through Open Router for developers working with limited budgets

2

.

Reinforcement Learning at Scale Drives Performance Gains

The breakthrough behind MiMo-V2.6 stems from scaled RL training across verifiable and complex tasks

3

. Xiaomi streamed the production run live, completing 30 RL steps across roughly 750,000 trajectories in under six days

3

. Training costs reached approximately $0.85 million for Flash and $2.62 million for Pro

3

. The average pass rate increased by 25% for Flash and 12% for Pro in relative terms. On DeepSWE v1.1, a held-out long-horizon software engineering benchmark, scores jumped from 48.8 to 65.68 for Flash and from 58.4 to 72.57 for Pro

3

. The company scaled RL compute through larger batches using a fully asynchronous architecture with 1,568 samples per update, processing 3.5-3.7 billion tokens per step

3

.

Building an Open Agent Stack Beyond Foundation Models

Xiaomi's release represents another step toward building a complete open agent stack encompassing foundation models, coding agents, harnesses, and training infrastructure

1

. The company introduced MiMo Code in June as an open-source terminal coding agent with persistent cross-session memory and task checkpoints. Internal testing suggested MiMo Code's advantage over Claude Code became more pronounced after workflows stretched beyond 200 execution steps

1

. Xiaomi also released HarnessX, a research framework treating prompts, memory systems, tools, and control logic as components that can be rewritten and optimized. The company reported an average 14.5% absolute performance gain across 15 model-benchmark combinations when the harness evolved dynamically

1

. Xiaomi is open-sourcing the technical report, training environments, and RL code for reproduction and verification

3

.

Multimodal AI Capabilities Across Diverse Applications

Both models are natively omnimodal, supporting a 1 million token context window with text, image, audio, and video input

1

2

. The models excel in 3D spatial reasoning, generating detailed and accurate 3D environments for game development, virtual reality, and simulation projects

2

. For robotics control, MiMo-V2.6-Pro enhances precision and adaptability in autonomous systems, supporting applications in manufacturing and logistics

2

. The model can take an image, video, or text prompt, break it into tasks, and coordinate multiple agents to build 3D scenes, perform visual verification, and produce runnable interactive worlds

3

. It can generate 3D objects and scenes in Blender from text descriptions or reference images for animation, 3D printing, and game development

3

.

Pro Ultra Speed Mode and Agentic Models for Production Workloads

Xiaomi is releasing MiMo-V2.6-Pro-UltraSpeed, which generates outputs up to 20 times faster than the standard Pro version

1

2

. This Pro Ultra Speed mode significantly reduces processing times for time-sensitive workflows. The architecture builds on the sparse mixture-of-experts approach established in MiMo-V2.5-Pro, a 1.02-trillion-parameter model with 42 billion parameters active during inference

1

. These agentic models were trained specifically for long-horizon software engineering with harness awareness—the ability to manage memory and context while operating inside agent scaffolds over hundreds or thousands of tool calls

1

. The Flash variant targets high-volume production workloads while retaining the same 1 million token context window and native multimodal capabilities

1

.

Defenses Against Reward Hacking in Scaled RL Training

As training scaled, Xiaomi froze the router to suppress training drift and implemented defenses against reward hacking

3

. These measures include reward design, adversarial evaluation, anomaly detection, and cross-checking between verifiers to improve training stability and reward reliability

3

. For agentic RL across different tasks, Xiaomi developed a unified trajectory representation and penalty mechanism, high-concurrency interaction across agent frameworks, a decoupled control and data plane, stabilized per-task sampling ratios, and optimized training and inference engines

3

. The training suite covers coding, general agents, visual, and cyber tasks across several harnesses

3

. Watch how Xiaomi's approach to grader compute provides more precise and diverse reward signals for long-horizon RL tasks, helping reduce paths and tokens per task while maintaining training effectiveness across the expanding model family.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved