Xiaomi-Robotics-1: scaling vision-language-action models with 100K+ hours of real-world trajectories
Xiaomi Robotics
Xiaomi Robotics presents a vision-language-action foundation model trained on more than 100,000 hours of real-world manipulation data collected via UMI devices, using a two-stage pretrain/post-train pipeline with auto-generated language annotations. The model sets state-of-the-art results on RoboCasa365 (57.6% success) and RoboDojo benchmarks.
Why it matters
Trending on Hugging Face Daily Papers with 72 upvotes, one of the most-discussed robotics papers of the week and a data point on how much real-world trajectory data is needed to scale VLA models.
Importance: 3/5
Notable research release with SOTA robotics benchmark results and strong community attention (72 HF Daily upvotes, below the 100-upvote bump threshold).
Sources
official
Xiaomi-Robotics-1 on HF Daily Papers
official
Xiaomi-Robotics-1 arXiv preprint