Score
Designs, implements, and analyzes algorithms and end-to-end training pipelines that learn action policies from demonstrations and expert trajectories—covering behavioral cloning and scalable supervised/optimization-based approaches as well as advantage-weighted, chunk- or warp-based weighting schemes and few-shot adaptation. Builds hierarchical and goal-conditioned imitation policies, methods for handling mixed-quality and teleoperated demonstrations, procedures for distilling policies into interpretable models, and evaluation protocols and metrics for measuring imitation performance and trajectory diversity.
Current imitation learning faces core challenges including poor generalization, severe covariate shift, heterogeneous expert data modalities (e.g., partially observable or unlabeled sequences), and the absence of a taxonomy tailored to the deep learning era. This paper presents a systematic survey of deep imitation learning advances since 2015. We propose a novel four-dimensional classification framework centered on expert data modality—namely, state-action pairs, trajectories, observation sequences, and unlabeled sequences—moving beyond traditional paradigms (behavioral cloning, inverse reinforcement learning, adversarial imitation). We critically analyze mainstream methods regarding theoretical assumptions, robustness, and empirical evaluation practices. Integrating recent technical developments, we identify key bottlenecks and articulate concrete future directions. The work provides systematic theoretical foundations and practical guidelines for algorithm design, benchmark construction, and cross-task transfer.
Demonstrating high-quality, physically plausible trajectories for teleoperated dexterous manipulation in contact-rich environments remains challenging due to the difficulty of acquiring consistent, diverse, and kinematically feasible human demonstrations. Method: This paper proposes a model-driven trajectory generation framework. It first identifies the high-entropy, low-consistency behavior of sampling-based planners (e.g., RRT) in contact-rich settings; then introduces a three-stage pipeline—RRT initialization, MPC-based refinement, and diffusion-model-based resampling—to jointly ensure physical feasibility, consistency, and diversity. Furthermore, it develops a goal-conditioned diffusion behavioral cloning (DBC) policy. Results: The method achieves zero-shot hardware transfer on two challenging contact-rich manipulation tasks, outperforming conventional behavioral cloning and pure planning baselines in terms of success rate, robustness, and generalization—without requiring any real-world demonstration data.
This paper identifies a fundamental error-amplification problem in imitation learning for continuous state-action spaces: even under stable system dynamics and smooth, deterministic expert policies, any smooth deterministic imitator incurs execution errors that grow exponentially with task horizon—a theoretical bottleneck pervasive in behavioral cloning and offline reinforcement learning. The authors provide the first rigorous proof of this exponential amplification in continuous-action settings, and identify viable mitigation strategies: adopting nonsmooth, non-Markovian, or highly stochastic policies, or leveraging expert datasets with sufficient action-space dispersion. Leveraging contraction theory and control-theoretic analysis, the paper further demonstrates that parametrization techniques—such as action chunking and diffusion-based policy representations—significantly suppress error accumulation. These results establish critical theoretical limits for robotic imitation learning and yield novel design principles for robust policy learning in continuous domains.
This work addresses the challenge that existing behavior cloning methods rely on first-person aligned data and struggle to learn effective policies from third-person passive observations. The authors propose a Mirror Learning framework that, for the first time, integrates viewpoint transformation with inverse dynamics modeling. By fine-tuning a video diffusion model to translate third-person observations into first-person perspectives and employing an inverse dynamics model to infer action trajectories, the method generates pseudo-first-person expert demonstrations from purely observational videos. This approach constructs a generative world model capable of training high-performance policies using only mirrored data, substantially reducing reliance on teleoperated demonstrations. When combined with first-person behavior cloning, the framework further enhances downstream policy performance.
Existing behavior cloning research lacks open and reproducible infrastructure, resulting in high barriers to policy development and hindering fair comparisons. This work proposes ABC—an end-to-end open-source behavior cloning framework that includes the large-scale teleoperated dataset ABC-130K (comprising 400 hours of simulated and real-world data), open-source hardware designs, a training framework, and a simulation pipeline. The framework innovatively integrates sim-to-real co-training and joint evaluation mechanisms. Leveraging the Diffusion Transformer (DiT) and vision-language-action (VLA) architectures, policies trained within this ecosystem demonstrate strong performance and generalization on dexterous manipulation tasks such as folding cardboard boxes and retrieving cards from wallets, thereby validating the effectiveness of the proposed infrastructure.
This paper addresses the high cost and poor generalization of imitation learning due to its reliance on high-quality expert demonstrations. To mitigate this dependency, we propose a meta-learning framework tailored for suboptimal demonstrations. Our method integrates weighted behavioral cloning with explicit policy distance regularization. Specifically, (1) we design the first meta-learned action ranker that dynamically reweights non-expert demonstrations using an advantage function; and (2) we introduce a learnable meta-objective that explicitly constrains the learned policy’s divergence from the expert policy. Evaluated on multi-task benchmarks, our approach significantly outperforms existing methods for learning from suboptimal demonstrations, achieving higher demonstration efficiency and improved policy performance while substantially reducing reliance on expert data.
This work addresses the high cost and suboptimality of acquiring high-quality demonstration data in imitation learning, particularly the lack of scalable data sources for goal-conditioned control tasks. To overcome this challenge, the authors propose an efficient data generation and augmentation framework that leverages trajectory optimization to automatically produce thousands of near-optimal trajectories within minutes on a standard laptop. By relabeling intermediate states along these trajectories as new goals, the training dataset is expanded by an order of magnitude. A lightweight goal-conditioned policy trained on this augmented dataset—containing fewer than 80,000 parameters—achieves near-optimal performance and high success rates across multiple tasks. Moreover, its inference speed exceeds that of the trajectory optimization solver by over 6,000×, substantially improving generalization and enabling practical deployment on embedded systems.
Acquiring expert demonstration data on real robots is prohibitively expensive, limiting the generalization and robustness of behavior cloning. To address this challenge, this work proposes ExpertGen, a framework that leverages imperfect priors—such as human or large language model demonstrations—in simulation to enable efficient and safe policy transfer. By freezing a pretrained diffusion policy and optimizing only its initial noise, ExpertGen integrates diffusion models, reinforcement learning, and DAgger to significantly improve task success rates under sparse rewards without requiring reward engineering, while preserving behavioral similarity to human demonstrations. Experimental results demonstrate that the method achieves success rates of 90.5% in industrial assembly tasks and 85% in long-horizon manipulation tasks, substantially outperforming baseline approaches, and successfully transfers to real-world robotic deployment.
Imitation learning is a popular paradigm to teach robots new tasks, but collecting robot demonstrations through teleoperation or kinesthetic teaching is tedious and time-consuming. In contrast, directly demonstrating a task using our human embodiment is much easier and data is available in abundance, yet transfer to the robot can be non-trivial. In this work, we propose Real2Gen to train a manipulation policy from a single human demonstration. Real2Gen extracts required information from the demonstration and transfers it to a simulation environment, where a programmable expert agent can demonstrate the task arbitrarily many times, generating an unlimited amount of data to train a flow matching policy. We evaluate Real2Gen on human demonstrations from three different real-world tasks and compare it to a recent baseline. Real2Gen shows an average increase in the success rate of 26.6% and better generalization of the trained policy due to the abundance and diversity of training data. We further deploy our purely simulation-trained policy zero-shot in the real world. We make the data, code, and trained models publicly available at real2gen.cs.uni-freiburg.de.
Behavioral cloning suffers from error accumulation during deployment due to out-of-distribution states and limited generalization. This work proposes a semi-parametric, retrieval-based imitation learning approach that reparameterizes policy modeling as a local neighborhood structure: by retrieving the k nearest neighbor states from expert demonstrations along with their corresponding actions, and predicting actions using relative distance vectors between states. The method requires no additional data, online feedback, or task-specific priors. Evaluated on continuous control and robotic manipulation tasks, it substantially outperforms standard behavioral cloning, achieving performance gains of 15%–46%, and is compatible with diverse state representations, including high-dimensional visual inputs.
This work addresses the limitations of traditional reinforcement learning in policy reuse and the restricted applicability of motion imitation to fixed trajectories. The authors propose a three-stage framework: first, an expert policy is trained to imitate human motion; second, policy distillation yields a frozen Hybrid Motion Prior (HMP) comprising a proprioceptive encoder, a residual vector quantization (RVQ) codebook, and an action decoder; third, downstream tasks reuse the HMP by selecting discrete codebook entries. The key innovations include the first formulation of imitated skills as a shareable, frozen prior, the interpretability of RVQ codebook entries—such as modulating gait via activated layers—and the introduction of rotation-based techniques to refine latent space structure and reduce falls. Experiments demonstrate that HMP significantly enhances training efficiency and stability in speed tracking, navigation, and fall recovery tasks, with successful real-world deployment on the Unitree G1 robot.