🤖 AI Summary
This study addresses the challenge of missing supervision in robot imitation learning under constrained interaction, caused by overlooking discrepancies between instructions and states. It reveals, for the first time, the local validity of instruction information in such scenarios and proposes an annotation-free, architecture-agnostic instruction–state discrepancy weighting mechanism. By integrating task phase analysis with a selective instruction retention strategy, this approach transforms response latency, progress, and demand variations into continuous dynamic supervision weights. Empirical evaluations demonstrate that the proposed method significantly outperforms uniform instruction supervision baselines on real-robot constrained tasks while achieving comparable performance on unconstrained tasks, thereby validating its effectiveness and generalizability.
📝 Abstract
Constructing action targets from measured robot motion is an established approach in imitation learning. Under interaction constraints, however, command-state discrepancy may reflect control demands that motion alone does not capture. We investigate when this information matters and how to exploit it. Across three real-robot tasks, task and phase analyses reveal larger supervision gaps under constrained interaction, while selective command retention provides evidence of locally useful command information. Building on these findings, we propose Command-State Discrepancy Weighting (CSDW), which accounts for robot response times and combines subsequent progress, persistent unmet demand, and demand changes into continuous weights for command supervision. The method requires no task-phase annotations or changes to policy architecture or inference. CSDW improves over uniform command supervision on constrained tasks, while methods perform similarly in the less constrained task.