Qwen team publishes Qwen-Drive-1.0, a vision-language foundation model for autonomous driving
Alibaba
Published August 31 (arXiv 2609.00111) by Qwen/Alibaba and HUST authors, Qwen-Drive-1.0 extends a pretrained VLM backbone with a bird's-eye-view perception head doing joint 3D detection, semantic occupancy and map segmentation, plus a Planning Expert generating ego trajectories. A staged training recipe mixes driving supervision with general vision-language data; evaluations span open-loop, pseudo-closed-loop and closed-loop settings. No weights released yet; featured on Hugging Face Daily Papers September 2.
Why it matters
First Qwen-branded driving foundation model, paper-only so far
Importance: 2/5
Notable domain foundation-model paper from the Qwen team