Beyond State-as-Action: Exploiting Command-State Discrepancy for Robot Imitation Learning

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of missing supervision in robot imitation learning under constrained interaction, caused by overlooking discrepancies between instructions and states. It reveals, for the first time, the local validity of instruction information in such scenarios and proposes an annotation-free, architecture-agnostic instruction–state discrepancy weighting mechanism. By integrating task phase analysis with a selective instruction retention strategy, this approach transforms response latency, progress, and demand variations into continuous dynamic supervision weights. Empirical evaluations demonstrate that the proposed method significantly outperforms uniform instruction supervision baselines on real-robot constrained tasks while achieving comparable performance on unconstrained tasks, thereby validating its effectiveness and generalizability.
📝 Abstract
Constructing action targets from measured robot motion is an established approach in imitation learning. Under interaction constraints, however, command-state discrepancy may reflect control demands that motion alone does not capture. We investigate when this information matters and how to exploit it. Across three real-robot tasks, task and phase analyses reveal larger supervision gaps under constrained interaction, while selective command retention provides evidence of locally useful command information. Building on these findings, we propose Command-State Discrepancy Weighting (CSDW), which accounts for robot response times and combines subsequent progress, persistent unmet demand, and demand changes into continuous weights for command supervision. The method requires no task-phase annotations or changes to policy architecture or inference. CSDW improves over uniform command supervision on constrained tasks, while methods perform similarly in the less constrained task.
Problem

Research questions and friction points this paper is trying to address.

Imitation Learning
Command-State Discrepancy
Interaction Constraints
Robot Control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Imitation Learning
Command-State Discrepancy
CSDW
Robot Manipulation
Supervision Weighting
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Peiyan Li
Peiyan Li
Ludwig-Maximilians-Universität München
data mininggraph mining
Y
Yueran Tao
Tsinghua University; SEEN·E Robotics
Enhao Zhang
Enhao Zhang
University of Washington
databasesvideo analyticshuman-computer interaction
Z
Zhixuan Zhao
Tsinghua University; SEEN·E Robotics
C
Chenghao Yue
Tsinghua University; SEEN·E Robotics
H
Hao Wang
Dalian University of Technology; SEEN·E Robotics
L
Lei Lv
Tongji University; SEEN·E Robotics
W
Wentao Zhao
Tsinghua University; SEEN·E Robotics
J
Jiahao Chen
Peking University; SEEN·E Robotics
X
Xin Liu
Tsinghua University; SEEN·E Robotics
Kangyao Huang
Kangyao Huang
Tsinghua University
Robot LearningAerial Robotics
Y
Yu Luo
Tsinghua University; SEEN·E Robotics
Huaping Liu
Huaping Liu
Professor of Electrical Engineering, Oregon State University
Communication theorywireless communicationssignal processingsensor networksinformation security