How does a World Action Model work?
Instead of mapping an image and text instruction directly to a motor command, a WAM learns physical transitions from video and action data. NVIDIA describes a joint video-action diffusion transformer that predicts latent visual representations of future states together with actions.
WAM vs VLA vs world model
| Model | Main task |
|---|---|
| World model | Predict how an environment evolves. |
| Vision-Language-Action (VLA) | Map visual input and language instructions to actions. |
| World Action Model (WAM) | Jointly predict world changes and the robot actions associated with them. |
Why does WAM matter?
Jointly learning dynamics and action could help robots transfer skills to unfamiliar tasks and environments. It may also reduce dependence on costly robot demonstrations by using large-scale human and internet video. It does not eliminate the need for embodiment-specific adaptation or solve all 3D spatial-reasoning problems.
DreamZero and NVIDIA GR00T N2
NVIDIA describes GR00T N2, based on DreamZero research, as a next-generation robot foundation model using a world action model architecture. NVIDIA reports more than twice the success rate of leading VLA baselines on new tasks and environments in its evaluations; broader independent validation remains important. Availability is planned by the end of 2026.
Applications and related terms
WAMs are relevant to manipulation, zero-shot generalization, whole-body control, human-to-robot transfer and Physical AI. Automotive autonomy also uses world models and action reasoning, but not every autonomous-driving system is a WAM.
Sources
NVIDIA: What is a World Action Model?; NVIDIA GR00T N2 announcement.
Automotive applications and further reading
Read about Physical AI in vehicles and XPENG G9L's VLA 2.0 technology. WAM is a research direction, not a claim that XPENG G9L itself uses NVIDIA GR00T N2.
Automotive examples and related reading
Explore Physical AI in vehicles and XPENG G9L and VLA 2.0. This does not imply that G9L uses NVIDIA GR00T N2.
Robot learning from demonstrations
Explore robot in-context learning and physical self-play alongside WAM and VLA.