Qwen-Drive-1.0: Alibaba open-sources a 4B vision-language foundation model for autonomous driving
On 2026-09-03 Alibaba's Qwen team released Qwen-Drive-1.0 (open weights, 4B, built on Qwen3.5-4B). Qwen calls it the first vision-language foundation model for autonomous driving that unifies 3D perception and visual QA at pretraining and extends to motion planning. The pretrained VLM is left unchanged; a BEV perception head and a flow-matching Planning Expert are attached as separate modules.
Key facts
- Base: Qwen3.5-4B; weights Qwen/Qwen-Drive-1.0-4B on Hugging Face (repo created 2026-08-27; blog 2026-09-03)
- External BEV head does 3D detection, semantic occupancy and BEV map segmentation
- Planning Expert: a diffusion transformer that generates 5-second ego trajectories via flow matching
- Driving QA average 69.43 (SFT model), ahead of the general and driving VLMs Qwen compared, incl. Alpamayo-1.5-10B, Cosmos-Reason2-8B and MiMo-Embodied-7B
- Examples: LingoQA 77.8, WaymoQA All 74.47; general scores stay close to the base (MMMU 72.67 vs 73.44)
What happened
Qwen adapted its small multimodal model for driving without changing its architecture. It trained in stages on combined public driving datasets plus general vision-language data to avoid forgetting, and attached two external modules for 3D perception and trajectory planning. Qwen reports competitive open-loop, pseudo-closed-loop and closed-loop planning results.
Why it matters
It is an open, small base model for teams building driving VLAs, competing with NVIDIA's Alpamayo and Cosmos models and Xiaomi's MiMo-Embodied. It also extends Qwen's robotics push (the Qwen-Robot suite, June 2026) to vehicles.
All benchmark numbers are Qwen's own.
Changelog
- 2026-09-30: created
Sources (4)
- officialQwen - Qwen-Drive-1.0
- paperarXiv 2609.00111 - Qwen-Drive-1.0
- codeHugging Face - Qwen/Qwen-Drive-1.0-4B
- codeGitHub - QwenLM/Qwen-Drive-1.0
id: 2026-09-03-qwen-drive-1-0 · updated 2026-09-30 · open in the interactive timeline