Updated October 2, 2026

Robot In-Context Learning: Can Robots Learn a New Task From One Video?

What robot in-context learning means, how video demonstrations guide unseen tasks, how it differs from fine-tuning, and where Skild AI's S1 and physical self-play fit.

Definition: Robot in-context learning is the ability to use a demonstration supplied at inference time to perform a new task without necessarily updating the robot model's weights.

How can one video teach a robot?

A video may demonstrate a goal, sequence of actions or desired final state. A trained robot foundation model interprets that example alongside live camera input and uses its learned physical representations to generate actions. The claim is task-dependent: one-video performance in a demonstration does not imply that any robot can learn any task from any clip.

In-context learning versus fine-tuning

MethodWhat changes?
In-context learningA demonstration is provided as context; model weights can remain unchanged.
Fine-tuningModel weights are updated using additional task-specific training.
Imitation learningA broader training approach that learns from demonstrations; it may use either training-time data or contextual demonstrations.

Skild AI S1 and single-video learning

Skild AI's S1 work has been presented as a step toward executing previously unseen long-horizon tasks from a single video demonstration without task-specific weight updates. Treat company demonstrations as reported research outcomes, not proof of general reliability across every robot, task or environment.

What is physical self-play?

Physical self-play applies self-improvement ideas from game AI to robot skill learning. After broad pretraining, a robot policy can practice through simulated interaction and reinforcement learning, discovering stronger strategies without requiring a human to demonstrate every variation. It is complementary to in-context learning rather than synonymous with it.

Why this matters for Physical AI

Robots deployed in warehouses, manufacturing and potentially vehicle assembly need to adapt to changing parts, layouts and tasks. In-context learning could shorten setup time, while physical self-play could improve skills through practice. Challenges include safety, distribution shift, embodiment differences and long-horizon error accumulation.

How it relates to VLA and world action models

A VLA maps vision and language to actions; in-context learning describes a way to adapt behavior using examples. A World Action Model jointly models future states and actions. These are different, potentially complementary concepts. Also see zero-shot generalization, whole-body control and Physical AI in automotive manufacturing.

Research and further reading

Skild AI official research and announcements · NVIDIA Research. This article will be updated with independently reproduced benchmarks as they become available.