Submitted:
08 September 2026
Posted:
09 September 2026
You are already at the latest version
Abstract
A single demonstration can guide a robot toward a new goal, execution sequence, or object correspondence without updating its learned parameters. Yet “one demonstration” describes neither a common information budget nor a common learning problem. This survey examines demonstration-driven in-context learning for robotic manipulation through four axes: context information, inference target, adaptation mechanism, and evaluated transfer. It connects early conditional policies to geometric methods, retrieval, vision-language-action models, and world action models. Three findings organize the comparison. Geometric structure reduces the transformation that a policy must learn, while making perception and correspondence more consequential. Compressed task representations favor reusable intent, whereas detailed context supports stage-dependent control at greater inference cost. Human-video adaptation obtains executable correspondence from prior pairing, motion estimation, or robot-future models; the video alone does not provide the same information in each route. Representative method tables make these preparation requirements explicit, and benchmark evidence locates current strengths in task selection and structured transfer. The resulting research questions concern informative example selection, selective access during execution, and interaction that resolves correspondence uncertainty.
Keywords:
embodied artificial intelligence
; in-context learning
; one-shot imitation learning
; robot learning
; vision-language-action models
; world action models
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.