#robotics
- NVIDIA Releases Cosmos 3: Open Omnimodal World Foundation Model for Physical AI NVIDIA research
- Kairos: A Native World Model Stack for Physical AI ACE Robotics research
- Black Forest Labs launches FLUX 3, a multimodal model generating video with synced audio Black Forest Labs video
- HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone Simple AI (Simple World Lab) research
- Google DeepMind unveils Gemini Robotics 2 for whole-body humanoid control Google DeepMind research
- TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM research
- AI2 Open-Sources MolmoAct2: Robotics VLA That Claims to Beat GPT-5 on Embodied Reasoning AI2 research
- Humanoid-GPT: Scaling to 2B Motion Frames Enables Zero-Shot Generalization in Humanoid Control research
- Alibaba Releases Qwen-RobotSuite: Three Embodied AI Foundation Models Alibaba / Qwen models-llm
- Alibaba Launches Qwen-Robot Suite: Three Foundation Models for Embodied AI and Robotics Alibaba / Qwen models-llm
- ENPIRE: AI Coding Agents Close the Loop on Physical Robotics Research Without Human Intervention NVIDIA / Carnegie Mellon University / UC Berkeley research
- World Action Models: A Survey National University of Singapore research
- Yandex Self-Driving Truck Completes First Fully Autonomous 700km Moscow–Saint Petersburg Run Yandex industry
- GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation GigaAI research
- LingBot-VLA 2.0: Bridging the Gap Between Foundation VLA Models and Real-World Deployment LingBot Team research
- Mistral Releases Robostral Navigate: Single-Camera Robot Navigation Model Mistral research
- InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization InternRobotics research
- RLDX-1: Multi-Stream Action Transformer Achieves 86.8% on ALLEX Humanoid Tasks RLWRLD research
- PhysBrain 1.0: Human Egocentric Video as Robot Training Data for VLA Models (133 HF upvotes) DeepCybo research
- ABot-AgentOS: General Robotic Agent OS with Lifelong Multi-modal Memory Alibaba research
- ABot-N1: Visual Language Navigation Foundation Model with Slow-Fast Architecture Alibaba research
- BadWAM: When World-Action Models Dream Right but Act Wrong research
- Xiaomi-Robotics-1: scaling vision-language-action models with 100K+ hours of real-world trajectories Xiaomi Robotics research
- Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Xiaomi Robotics research
- RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model Alibaba DAMO Academy research
- Anthropic's Project Pilot tests whether AI models can fly drones Anthropic research
- World Action Models: First Systematic Survey of Embodied Foundation Models Unifying World Modeling and Action OpenMOSS research
- Hallucination in World Models is Predictable and Preventable UC San Diego research
- Tencent releases HY-Embodied-0.5-X update for embodied agents Tencent models-llm
- Playful Agentic Robot Learning: Self-Directed Play Yields Transferable Robot Skills UC Berkeley research
- PhysisForcing: Physics-Reinforced World Models Improve Robot Manipulation Success by 50% Peking University / NVIDIA research