Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
Xiaomi Robotics
Xiaomi Robotics presents a vision-language-action foundation model pretrained on over 100,000 hours of real-world manipulation trajectories using an automatic natural-language scene-change annotation pipeline, then aligned to specific robot platforms. It achieves 57.6% success on RoboCasa365 (vs. prior best 46.6%) and a score of 20.07 on RoboDojo (vs. prior best 13.07), with code and weights promised for release.
Why it matters
This paper has 215 upvotes on HuggingFace Daily Papers, the highest in the period, indicating strong community interest in large-scale real-world VLA data scaling.
Importance: 3/5
HF Daily Papers top-voted (215 upvotes, >=100 threshold) -> +1 bump from base 2.