ReShoot: Generative Visual Domain Randomization of Recorded Robot Demonstrations for Visuomotor Policy Learning

📅 2026-09-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决机器人模仿学习策略在视觉条件变化下的性能下降问题,ReShoot通过重新渲染已录制的演示视频来生成视觉多样性,提高策略对不同视觉环境的鲁棒性。
📝 Abstract
Imitation-learned robot policies are frequently overfit to the visual conditions present in their training demonstrations. Consequently, variations in object color or background appearance often induce substantial performance degradation. A common mitigation strategy is to acquire additional demonstrations in each novel visual context; however, this approach is resource-intensive, requiring repeated access to a robot, a controlled environment, and human operation for every appearance condition to be covered. We introduce ReShoot, a framework that synthesizes visual diversity by re-rendering previously recorded demonstrations under altered appearances, thereby shifting the burden from data collection to generation. A vision-language model captions the scene, edits a targeted attribute (e.g., background, object color, or material), and an edge-conditioned video generator re-renders both camera views to match. The instruction is updated accordingly. The action sequence and proprioceptive trajectory are copied verbatim without relabeling, so each generated episode retains the recorded action and proprioceptive labels. On LIBERO, a policy trained on an equal mixture of recorded and re-rendered demonstrations matches the performance of recorded-only training (96.5% vs. 96.9%). Moreover, the mixed training set improves robustness to scene perturbations on LIBERO-Plus (85.5% vs. 82.3%). Across two physical robotic platforms, deploying ReShoot with 43 and 100 pre-collected demonstrations increased the success rate on recolored objects from 0.0% to 42.9% and 47.5%, respectively, while maintaining performance under the original recorded appearance.
Problem

Research questions and friction points this paper is trying to address.

imitation learning
visual conditions
performance degradation
data collection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generative Visual Domain Randomization
Imitation Learning
Visuomotor Policy Learning
Scene Perturbation Robustness
🔎 Similar Papers
No similar papers found.
C
Chiyoung Kim
Chung-Ang University, Seoul, Republic of Korea
M
Min Sung Choi
Chung-Ang University, Seoul, Republic of Korea
J
Jinho Ju
Chung-Ang University, Seoul, Republic of Korea
C
Chanhoe Gu
Chung-Ang University, Seoul, Republic of Korea
D
Donghwan Hwang
Chung-Ang University, Seoul, Republic of Korea
Wonseok Choi
Wonseok Choi
PhD Student, POSTECH
vision language modelmodel evaluationcomputer vision
W
Woongsun Jeon
Chung-Ang University, Seoul, Republic of Korea
Minhyeok Lee
Minhyeok Lee
Yonsei University
Computer Vision and Pattern Recognition