Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

📅 2026-07-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work proposes a general-purpose robotic foundation model designed to understand and execute diverse language instructions while achieving strong generalization in unseen environments. Built upon a vision-language-action (VLA) architecture, the approach employs a two-stage training paradigm: first, pretraining on over 100,000 hours of real-world manipulation trajectories, followed by fine-tuning using an automated annotation pipeline that generates state-change descriptions aligned with high-level instructions. The resulting model achieves a state-of-the-art success rate of 57.6% on RoboCasa365 and an average score of 20.07 on RoboDojo, substantially outperforming existing methods. These results demonstrate exceptional zero-shot transfer capability and data-efficient adaptation across varied tasks and environments.
📝 Abstract
We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) efficiently adapting to novel downstream tasks with minimal fine-tuning data. We propose a two-stage training recipe consisting of pre-training and post-training. During pre-training, we imbue the model with broad and generalizable action-generation capabilities by training on over 100k hours of real-world manipulation trajectories collected via UMI devices. Crucially, we develop a scalable auto-labeling pipeline that annotates trajectory clips with natural languages describing scene state transitions, providing rich and precise conditioning for action learning. During post-training, we aim to align these capabilities with robot embodiments and imperative instructions that humans naturally use to prompt robots. Extensive experiments demonstrate strong scaling behavior. Xiaomi-Robotics-1 consistently improves with increased data scales and model sizes during pre-training. This scaling behavior directly transfers to post-training, where a stronger pre-training model yields better out-of-the-box real-robot performance in unseen environments. Furthermore, Xiaomi-Robotics-1 serves as a strong robot foundation policy that can be efficiently fine-tuned on complex, dexterous tasks with high data efficiency. Across multiple simulation benchmarks, Xiaomi-Robotics-1 outperforms state-of-the-art methods. Notably, it establishes a new state-of-the-art with a 57.6% success rate on RoboCasa365, surpassing the previous best of 46.6%. Furthermore, it achieves an average score of 20.07 on RoboDojo, significantly outperforming the prior state-of-the-art (13.07). Code and model checkpoints will be released. Project page: https://robotics.xiaomi.com/xiaomi-robotics-1.html
Problem

Research questions and friction points this paper is trying to address.

vision-language-action
robotics
real-world trajectories
instruction following
foundation model
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action Model
Real-World Trajectories
Auto-Labeling Pipeline
Two-Stage Training
Robot Foundation Model
🔎 Similar Papers
No similar papers found.