In-context Robot Learning Made Simple: A Democratized Recipe for Manipulation Tasks

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the ambiguity of visual demonstrations and the absence of explicit task definitions in robotic in-context learning by formally defining this problem for the first time and proposing the SimpleICL framework. Methodologically, it introduces a visual prompt encoder alongside a low-cost data collection protocol to clarify learning objectives, enabling manipulation task inference without large-scale pretraining. Experiments demonstrate that the proposed framework achieves superior performance in both simulated and real-world environments, revealing critical discriminative properties such as action and semantic sensitivity. By fully open-sourcing the datasets and training pipelines, this work establishes a minimalist and reproducible research paradigm for robotic in-context learning.
πŸ“ Abstract
We study robotic in-context learning (ICL), an emerging paradigm that enables robots to infer and execute tasks from visual demonstrations. Despite its growing promise, the problem itself remains under-defined: a visual demonstration simultaneously conveys action trajectories, object semantics, manipulation affordances, spatial relations, and task goals, making it unclear what information the robot is actually expected to follow. In this work, we first provide a clear problem definition of robot ICL that explicitly defines its learning target and resolves this fundamental prompt ambiguity. Building on this definition, we develop a minimalist and reproducible ICL framework (SimpleICL) with a visual prompt encoder and a low-cost data collection protocol. Without massive pre-training or specialized data infrastructure, our framework achieves strong performance in both simulation and real-world environments. Extensive experiments further reveal several key properties of robot ICL, including action, semantic, composition, and affordance discrimination. We will fully open-source our data and training pipeline to facilitate systematic and reproducible research on robot ICL. The project page can be found at https://simpleicl.github.io/simpleicl.
Problem

Research questions and friction points this paper is trying to address.

robotic in-context learning
visual demonstration
prompt ambiguity
manipulation tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

In-context Learning
Robotic Manipulation
Visual Prompt Encoder
SimpleICL
Affordance Discrimination
πŸ”Ž Similar Papers
πŸ’Ό Related Jobs
No related jobs found.
M
Minxing Li
NLPR, Institute of Automation, Chinese Academy of Sciences (CASIA)
M
Minghao Han
NLPR, Institute of Automation, Chinese Academy of Sciences (CASIA)
W
Weizhi Zhao
NLPR, Institute of Automation, Chinese Academy of Sciences (CASIA)
Hanwen Wang
Hanwen Wang
Johns Hopkins University, SOM
Quantitative Systems PharmacologyOncologySystems Biology
X
Xiangshuo Liu
NLPR, Institute of Automation, Chinese Academy of Sciences (CASIA)
S
Shuyao Shang
NLPR, Institute of Automation, Chinese Academy of Sciences (CASIA)
J
Jingxiang Zhou
NLPR, Institute of Automation, Chinese Academy of Sciences (CASIA)
M
Mingchao Sun
Amap, Alibaba Group
Hongyu Pan
Hongyu Pan
Alibaba DAMO Academy, Autonomous Driving Lab
Computer VisionDetectionSegmentationPoint CloudMotion,End2End
Mu Xu
Mu Xu
alibaba
CV LLM VLM VLA RL
Yu Liu
Yu Liu
Alibaba Group
self-supervised learninggenerative modeling
L
Lue Fan
NLPR, Institute of Automation, Chinese Academy of Sciences (CASIA)
Zhaoxiang Zhang
Zhaoxiang Zhang
Institute of Automation, Chinese Academy of Sciences
Computer VisionPattern RecognitionBiologically-inspired Learning