Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. Qwen-Drive-1.0: Alibaba open-sources a 4B vision-language…

Qwen-Drive-1.0: Alibaba open-sources a 4B vision-language foundation model for autonomous driving

★★after cutoffroboticsAlibabaQwenconfidence: high

On 2026-09-03 Alibaba's Qwen team released Qwen-Drive-1.0 (open weights, 4B, built on Qwen3.5-4B). Qwen calls it the first vision-language foundation model for autonomous driving that unifies 3D perception and visual QA at pretraining and extends to motion planning. The pretrained VLM is left unchanged; a BEV perception head and a flow-matching Planning Expert are attached as separate modules.

Key facts

What happened

Qwen adapted its small multimodal model for driving without changing its architecture. It trained in stages on combined public driving datasets plus general vision-language data to avoid forgetting, and attached two external modules for 3D perception and trajectory planning. Qwen reports competitive open-loop, pseudo-closed-loop and closed-loop planning results.

Why it matters

It is an open, small base model for teams building driving VLAs, competing with NVIDIA's Alpamayo and Cosmos models and Xiaomi's MiMo-Embodied. It also extends Qwen's robotics push (the Qwen-Robot suite, June 2026) to vehicles.

All benchmark numbers are Qwen's own.

Changelog

  • 2026-09-30: created

Sources (4)

id: 2026-09-03-qwen-drive-1-0 · updated 2026-09-30 · open in the interactive timeline