Xiaomi-Robotics-1: scaling vision-language-action models with 100K+ hours of real-world trajectories

Xiaomi Robotics

Research official 2 src. ~1 min

Xiaomi Robotics presents a vision-language-action foundation model trained on more than 100,000 hours of real-world manipulation data collected via UMI devices, using a two-stage pretrain/post-train pipeline with auto-generated language annotations. The model sets state-of-the-art results on RoboCasa365 (57.6% success) and RoboDojo benchmarks.

Why it matters

Trending on Hugging Face Daily Papers with 72 upvotes, one of the most-discussed robotics papers of the week and a data point on how much real-world trajectory data is needed to scale VLA models.

Importance: 3/5

Notable research release with SOTA robotics benchmark results and strong community attention (72 HF Daily upvotes, below the 100-upvote bump threshold).

Sources