Score
Designs, builds, or evaluates policies by training supervised models to map observations or states to actions using recorded expert demonstrations, including the data collection, dataset preprocessing, model architecture, and training/evaluation pipeline required to learn from demonstrations. Analyzes common failure modes such as covariate shift and compounding errors and implements mitigation strategies (e.g., data augmentation, noise injection, or iterative dataset aggregation) to improve imitation performance.
Current imitation learning faces core challenges including poor generalization, severe covariate shift, heterogeneous expert data modalities (e.g., partially observable or unlabeled sequences), and the absence of a taxonomy tailored to the deep learning era. This paper presents a systematic survey of deep imitation learning advances since 2015. We propose a novel four-dimensional classification framework centered on expert data modality—namely, state-action pairs, trajectories, observation sequences, and unlabeled sequences—moving beyond traditional paradigms (behavioral cloning, inverse reinforcement learning, adversarial imitation). We critically analyze mainstream methods regarding theoretical assumptions, robustness, and empirical evaluation practices. Integrating recent technical developments, we identify key bottlenecks and articulate concrete future directions. The work provides systematic theoretical foundations and practical guidelines for algorithm design, benchmark construction, and cross-task transfer.
Imitation learning for multi-turn language model agents suffers from covariate shift: once the student policy deviates from the expert’s state-action distribution, it encounters out-of-distribution states unseen during training, leading to catastrophic generalization failure. To address this, we propose Online Expert Correction (OEC), a novel data-generation paradigm—“student-initiated, expert-intervened, trajectory-corrected”—that synergistically integrates offline demonstrations with online interaction to dynamically mitigate covariate shift. OEC performs end-to-end optimization via online rollouts, expert takeover upon detected divergence, reward-guided rejection sampling, and joint supervised fine-tuning. Evaluated on SWE-bench, OEC improves task success rates by 14% and 13% for 7B and 32B models, respectively, substantially overcoming the generalization bottleneck inherent in conventional offline imitation learning for sequential decision-making tasks.
In robot imitation learning, heterogeneous demonstration data—exhibiting varying quality and sparsity—degrades policy generalization: low-quality demonstrations are often imperceptible to human annotators yet significantly reduce test success rates. To address this, we propose Demo-SCORE, an autonomous demonstration filtering framework leveraging real-robot online interaction feedback. Its core innovation lies in using binary success/failure outcomes from real-world roll-outs as an unsupervised evaluation signal—first of its kind—combined with policy classifier training, cross-validation, and uncertainty-aware filtering to jointly optimize simulation-based pre-screening and real-world closed-loop feedback. Experiments demonstrate that policies trained on Demo-SCORE-filtered demonstrations achieve 15–35 percentage-point improvements in real-robot task success rates over full-dataset baselines, enabling automatic discovery and effective utilization of high-quality demonstrations.
Existing demonstration filtering metrics struggle to identify structural flaws that degrade imitation learning performance, particularly failing when relying solely on action information. This work constructs a controlled testbed by injecting known defects—such as action noise, tremor, truncation, and critical-step errors—to systematically evaluate the effectiveness of seven filtering metrics in detecting such flaws and improving downstream policy performance. The study reveals, for the first time, that action-only metrics are not only ineffective but actively harmful in the presence of structural errors, while state-aware metrics partially mitigate these issues yet recover at most one-third of the lost performance. Notably, high defect detection accuracy does not necessarily translate into policy gains. The authors open-source their testbed and metric implementations, underscoring the necessity of state-trajectory analysis in demonstration assessment.
This work addresses two core challenges in large-scale robotic manipulation datasets: (1) designing high-value diversity dimensions to enhance data utility, and (2) efficiently retrieving task-aligned demonstrations from existing datasets. To this end, we introduce a programmable data generation framework that explicitly models controllable diversity variables—including camera pose, object categories, and spatial layout. Our analysis reveals, for the first time, that camera pose and spatial arrangement are critical determinants of both dataset diversity and task alignment. We further propose a task-oriented demonstration retrieval algorithm grounded in geometric-semantic joint alignment. Evaluated on real-world datasets including DROID, our method improves downstream policy performance by up to 70%. Crucially, insights and gains observed in simulation generalize successfully to physical robot platforms, demonstrating robust cross-domain transferability.
To address the high cost and limited scale of high-quality real-world demonstration data in robotic imitation learning, this paper proposes a closed-loop learning framework integrating few-shot learning, simulation augmentation, and human-in-the-loop error correction. The method leverages a small set of real demonstrations (3–5 trials) augmented with synthetic data generated via GANs or diffusion models, followed by online policy fine-tuning to enable zero-shot cross-task transfer. Its key contributions are: (1) a novel vision-guided real-time human correction mechanism supporting natural multimodal interaction; (2) an end-to-end control pipeline built on ROS and PyTorch; and (3) empirical validation across manipulation tasks—including bottle collection, stacking, and hammering—achieving >92% success rates. Crucially, the framework transfers zero-shot to an unseen task—beverage tray arrangement—with 86% success, significantly outperforming pure-simulation baselines.
This study addresses the limitation of existing structured policy generation methods, which rely on manual or static knowledge and struggle to align with expert demonstrations. To overcome this, we propose a closed-loop iterative framework leveraging large language models (LLMs). The approach semanticizes rollout data into tabular formats, enabling LLMs to automatically diagnose and rectify structural deficiencies in policies. By utilizing tabularized rollout analysis as a feedback signal, the framework achieves automated alignment of policy structures without human intervention. Furthermore, this work integrates imitation learning with automated code generation techniques. Experimental results demonstrate that the proposed method improves performance by 15% while reducing computational costs by 75%, significantly enhancing overall policy generation efficiency.
为解决复杂任务中人类难以提供有效示例的问题,提出GLIDE框架,通过推断任务特定失败模式并生成指导规则来提高数据收集和策略执行的成功率。
This work addresses the challenge that existing behavior cloning methods rely on first-person aligned data and struggle to learn effective policies from third-person passive observations. The authors propose a Mirror Learning framework that, for the first time, integrates viewpoint transformation with inverse dynamics modeling. By fine-tuning a video diffusion model to translate third-person observations into first-person perspectives and employing an inverse dynamics model to infer action trajectories, the method generates pseudo-first-person expert demonstrations from purely observational videos. This approach constructs a generative world model capable of training high-performance policies using only mirrored data, substantially reducing reliance on teleoperated demonstrations. When combined with first-person behavior cloning, the framework further enhances downstream policy performance.
This work addresses the lack of a general, auditable dynamic control mechanism in existing training systems, which typically rely on framework-specific code. The authors propose the first cross-framework, open-source control plane that exposes training interfaces through a unified protocol, integrating declarative configuration, request validation, and secure control-point scheduling within the Aim workspace to enable metric monitoring, real-time intervention, and operational traceability. The system supports safe human and automated controller interventions during training while fully logging all operational trajectories. Experiments across five NLP and reinforcement learning tasks demonstrate its effectiveness, and the open-source implementation provides a foundation for reproducible human-in-the-loop training.
This work addresses the performance degradation of imitation learning policies caused by low-quality user demonstration data—such as jittery or oscillatory trajectories—in real-world scenarios. To this end, the authors propose an unsupervised, zero-interaction data quality assessment metric based on the power spectral density (PSD) of trajectories. This method requires neither policy training, environment interaction, nor expert annotations, enabling efficient ranking and selection of high-quality demonstrations for policy fine-tuning. As the first study to leverage PSD for demonstration filtering, it substantially reduces computational overhead. Extensive evaluations across multiple benchmarks and a user study with older adults demonstrate that policies trained on PSD-filtered data achieve significantly higher task success rates and smoother trajectories, outperforming both unfiltered baselines and alternative filtering approaches.