Physical AI · Robotics · Updated October 3, 2026

What Is an On-Device VLA?

An on-device vision-language-action model (VLA) interprets what a robot sees and the instructions it receives, then generates physical actions while running locally on robotic hardware instead of relying on a continuous cloud connection.

On-device VLA editorial illustration showing local inference
Original CarGlossary editorial illustration; conceptual depiction, not a photograph of a commercial robot.
In one sentence: An on-device VLA brings vision, language understanding and action generation onto the robot, reducing reliance on network connectivity.

How does an on-device VLA work?

Cameras and other sensors provide observations; a human supplies a natural-language task; and the VLA produces commands for the robot's supported joints, grippers or other actuators. Local inference means the core action model runs on the robot or its local compute hardware. Some deployments may still use remote services for planning, updates or monitoring.

On-device VLA vs cloud VLA

ConsiderationOn-device VLACloud-dependent VLA
ConnectivityCan continue local inference without a continuous internet connectionMay depend on network availability
LatencyAvoids cloud round trips, though local compute speed mattersCan use larger remote compute, but network delays vary
PrivacyCan keep raw sensor data local, depending on deploymentMay send observations to remote infrastructure
Compute constraintsLimited by onboard power, memory and thermal capacityCan access more powerful centralized hardware

Gemini Robotics On-Device 2: a current example

Google DeepMind describes Gemini Robotics On-Device 2 as its efficient VLA optimized to run locally on robotic devices. The company says the model is natively multi-embodiment and can adapt to completely new robot bodies with fewer than 200 examples gathered over a few hours. These are Google's reported results for its supported adaptation settings, not a guarantee for every robot or task.

As of October 3, 2026, Google lists On-Device 2 for trusted testers; its public waitlist is not equivalent to general model availability. Google's separate Gemini Robotics ER 2 embodied-reasoning model is available in public preview, but it is not the same product as On-Device 2.

What is multi-embodiment adaptation?

Different robots have different arms, joints, grippers, sensors and action spaces. A multi-embodiment model learns transferable representations across robot types. Adapting a model to a new robot may still require calibration, demonstrations and safety testing; it does not mean plug-and-play compatibility with every machine.

Does local inference make a robot safer?

Reduced network dependence can improve operational resilience, but an on-device VLA does not itself establish functional safety. Robots still need appropriate motion limits, emergency stops, human-proximity safeguards, validation and application-specific risk assessment.

Why does this matter to automotive manufacturing?

Factories often need reliable, low-latency robot behavior near moving equipment and intermittent network conditions. Local VLAs could support adaptable picking, material handling and assembly, although production deployment requires task-specific testing. Explore Physical AI in vehicles and manufacturing and whole-body control.

Related terminology

Sources and further reading

Related robot learning guides

Compare WAM vs VLA and explore robot in-context learning for inference-time adaptation.