Dynamics-Aware Meta-Imitation for Generalization to Unseen Robotic Manipulation

📅 2026-07-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limited generalization of robotic imitation learning in unseen tasks, which stems from data scarcity and environmental discrepancies. To overcome this, the authors propose the DAMI framework, which leverages meta-learning to construct a shared skill space and introduces three key components: a Visual-Motor Trajectory (VMT) module to model spatiotemporal dynamics, an Unpaired Unified Task (U2T) module to fuse multimodal observations without requiring paired demonstrations, and a Task-Conditioned Feature Modulation (TCFM) mechanism that emphasizes task-essential features over superficial cues. DAMI enables rapid adaptation to new tasks with only a few samples and no need for task-aligned demonstrations. Experimental results demonstrate that DAMI significantly outperforms existing methods in both simulation and real-world settings, achieving strong performance on seen tasks and exceptional generalization to unseen tasks after minimal fine-tuning.
📝 Abstract
Imitation Learning aims to learn skills from extensive observations and demonstrations for robots, so it suffers from data scarcity and environment generalization. The existing methods predominantly focus on imitation from in-domain tasks and consequently struggle with generalization to unseen tasks. To bridge this generalization gap, we propose the \textbf{D}ynamics-\textbf{A}ware \textbf{M}eta-\textbf{I}mitation (DAMI) framework. By integrating meta-learning to construct a shared skill space, DAMI equips agents for rapid adaptation to novel tasks. We introduce the Visual-Motor Trajectory (VMT) module to capture complex spatio-temporal dynamics within the task latent space. Furthermore, we propose the Unpaired Unified Task (U2T) block to fuse unstructured multimodal observations. To coordinate these representations, we integrate a Task-Conditioned Feature Modulation (TCFM) mechanism customized for modulating low-level 3D features. By capturing intrinsic dynamics from a random complete reference demonstration, our framework learns the underlying task logic rather than memorizing static cues, ensuring effective generalization. Extensive experiments in both simulation and real-world settings demonstrate that our approach outperforms state-of-the-art baselines regarding direct inference on seen tasks and adaptation to unseen tasks via few-shot fine-tuning.
Problem

Research questions and friction points this paper is trying to address.

Imitation Learning
Generalization
Robotic Manipulation
Meta-Learning
Unseen Tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

meta-imitation learning
dynamics-aware representation
visual-motor trajectory
task-conditioned feature modulation
unseen task generalization
🔎 Similar Papers
No similar papers found.