SynIL: Leveraging Synergy for Offline Imitation Learning from Imperfect Demonstration Datasets

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of offline imitation learning to suboptimal or noisy demonstrations, which typically necessitates manual pre-filtering. To overcome this limitation, we propose SynIL, a framework that pioneers the integration of motor synergy theory from neuroscience into offline reinforcement learning. By quantifying motor synergy manifestations to automatically assess demonstration quality, SynIL constructs transition-level dense reward signals without requiring expert reference data or task-specific heuristics, thereby enabling label-free policy learning. Evaluations on the D4RL and Robomimic benchmarks demonstrate that synergy-derived rewards exhibit high correlation with ground-truth rewards. Furthermore, the proposed method substantially outperforms behavior cloning and surpasses offline reinforcement learning algorithms based on true environment rewards in sparse-reward scenarios.
📝 Abstract
Imitation learning enables robots to acquire complex skills directly from massive demonstration datasets, but its performance degrades severely when datasets are contaminated with suboptimal or noisy demonstrations. While prior quality-assessment methods attempt to filter or reweight data, they typically rely on manual pre-selection of expert reference data or task-specific heuristics, limiting scalability. To address this challenge, we introduce SynIL (Synergy-based Imitation Learning), a novel framework for automated, label-free demonstration quality assessment in offline reinforcement learning. Grounded in neuroscientific evidence that motor synergy, a low-dimensional coordinated structure in movement, correlates directly with motor proficiency, SynIL algorithmically quantifies synergy manifestation to generate dense, transition-level reward signals via self-supervised reward regression. Comprehensive evaluations on D4RL locomotion benchmarks and multi-human Robomimic manipulation datasets demonstrate that synergy-derived rewards correlate strongly with ground-truth rewards. Furthermore, SynIL substantially outperforms Behavior Cloning (BC) and achieves performance comparable to, and in sparse-reward human teleoperation scenarios, superior to, offline reinforcement learning trained on true environment rewards.
Problem

Research questions and friction points this paper is trying to address.

Imitation Learning
Offline Reinforcement Learning
Imperfect Demonstrations
Demonstration Quality Assessment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Offline Imitation Learning
Motor Synergy
Self-supervised Reward Regression
Imperfect Demonstrations
Transition-level Rewards
🔎 Similar Papers
2024-06-12arXiv.orgCitations: 0