How does an on-device VLA work?
Cameras and other sensors provide observations; a human supplies a natural-language task; and the VLA produces commands for the robot's supported joints, grippers or other actuators. Local inference means the core action model runs on the robot or its local compute hardware. Some deployments may still use remote services for planning, updates or monitoring.
On-device VLA vs cloud VLA
| Consideration | On-device VLA | Cloud-dependent VLA |
|---|---|---|
| Connectivity | Can continue local inference without a continuous internet connection | May depend on network availability |
| Latency | Avoids cloud round trips, though local compute speed matters | Can use larger remote compute, but network delays vary |
| Privacy | Can keep raw sensor data local, depending on deployment | May send observations to remote infrastructure |
| Compute constraints | Limited by onboard power, memory and thermal capacity | Can access more powerful centralized hardware |
Gemini Robotics On-Device 2: a current example
Google DeepMind describes Gemini Robotics On-Device 2 as its efficient VLA optimized to run locally on robotic devices. The company says the model is natively multi-embodiment and can adapt to completely new robot bodies with fewer than 200 examples gathered over a few hours. These are Google's reported results for its supported adaptation settings, not a guarantee for every robot or task.
As of October 3, 2026, Google lists On-Device 2 for trusted testers; its public waitlist is not equivalent to general model availability. Google's separate Gemini Robotics ER 2 embodied-reasoning model is available in public preview, but it is not the same product as On-Device 2.
What is multi-embodiment adaptation?
Different robots have different arms, joints, grippers, sensors and action spaces. A multi-embodiment model learns transferable representations across robot types. Adapting a model to a new robot may still require calibration, demonstrations and safety testing; it does not mean plug-and-play compatibility with every machine.
Does local inference make a robot safer?
Reduced network dependence can improve operational resilience, but an on-device VLA does not itself establish functional safety. Robots still need appropriate motion limits, emergency stops, human-proximity safeguards, validation and application-specific risk assessment.
Why does this matter to automotive manufacturing?
Factories often need reliable, low-latency robot behavior near moving equipment and intermittent network conditions. Local VLAs could support adaptable picking, material handling and assembly, although production deployment requires task-specific testing. Explore Physical AI in vehicles and manufacturing and whole-body control.
Related terminology
- Vision-Language-Action (VLA) — the model category that connects perception, instructions and action.
- World Action Model (WAM) — jointly predicts world changes and actions; not synonymous with local inference.
- Robot in-context learning — adapting behavior from demonstrations without necessarily changing weights.
- Zero-shot generalization — performing unfamiliar tasks or operating in unseen environments.
Sources and further reading
- Google DeepMind: Gemini Robotics On-Device 2 — model overview and availability
- Google DeepMind: Gemini Robotics 2 announcement
- Google DeepMind: original On-Device announcement, June 2025
Related robot learning guides
Compare WAM vs VLA and explore robot in-context learning for inference-time adaptation.