Alibaba’s research team introduced Qwen-Drive 1.0, a unified driving model that combines 3D scene perception, traffic question answering, and short-horizon route planning while providing natural-language explanations for braking and steering decisions. Built on the Qwen3.5-4B vision-language model, Qwen-Drive 1.0 adds dedicated bird’s-eye-view perception and planning modules, and is trained on 24 public driving datasets plus custom explanation data to reduce catastrophic forgetting and strengthen spatial understanding. Benchmarks show improved performance over specialized systems, but explanations can misidentify hazards, and robustness to unfamiliar camera setups and potential adversarial attacks remains limited. The team released the model openly on Hugging Face, ModelScope, and GitHub.
This update represents a notable development in the Ai sector. Organizations and founders tracking this space should evaluate potential strategic and technical implications on their operations.