VLA means Vision-Language-Action: an AI model approach that combines visual perception, semantic or language-like reasoning, and action outputs for intelligent driving or other embodied systems.
How the VLA idea works
“Vision” represents perception of the physical environment. “Language” represents a richer semantic or reasoning layer rather than ordinary voice commands. “Action” means producing outputs that can influence behavior. The exact architecture varies by company, so VLA should be treated as a model family or design concept, not one standard.
Why VLA is appearing in car discussions
XPENG has made VLA prominent through its VLA 2.0 intelligent-driving system. In 2026 the company reported production rollout in China, public road testing for Robotaxi, and plans for broader international deployment. That makes VLA a buyer-facing term rather than only a research term.
VLA vs a world model
A VLA model maps perception and context toward actions. A world model tries to represent or predict how the environment may evolve. Manufacturers can combine the two approaches rather than choosing one exclusively.
Does VLA mean autonomous driving?
No. A VLA architecture does not by itself define an SAE automation level. Always separate the AI model used inside a system from the legally and operationally defined driving function delivered to the customer.
Related terms
Sources
Automotive terminology evolves quickly. We separate established standards from manufacturer-specific names and update pages as technologies move into production.