🤖 AI Summary
This study addresses the scarcity of robot data and the challenges of camera motion interference and temporal dynamic discrepancies in transferring first-person videos. To this end, this work proposes an interaction-centric spectral latent guidance framework. The method extracts interaction-centric latent actions by decoupling observer motion and aligns them through low-frequency shared components in the spectral domain to guide robot policy generation. Its core contribution lies in integrating latent action distillation, spectral analysis, and low-frequency component alignment to enable efficient cross-embodiment task semantic transfer. Experimental results demonstrate that the proposed framework achieves a 99.2% success rate on the LIBERO benchmark and exhibits strong performance in real-world multi-task scenarios.
📝 Abstract
Learning general-purpose robot policies requires large-scale real-world interaction data, yet collecting robot demonstrations remains expensive and difficult to scale. Egocentric videos offer abundant human interaction experience with task-relevant semantics for robotic manipulation, but direct transfer is challenging for two reasons: latent actions inferred from frame reconstruction can be dominated by nuisance variation such as ego-camera motion, and human and robot behaviors often exhibit different temporal dynamics. We propose WING (World Action Learning via INteraction-Centric Spectral Latent Guidance), a framework for transferring interaction knowledge from egocentric videos to robot policies. WING first separates observer-induced motion from hand-object interaction and distills the interaction-centric component into latent actions. It then exploits the observation that cross-embodiment task semantics are concentrated in slowly varying temporal structures, identifying shared low-frequency components between egocentric latent actions and robot behaviors in the spectral domain and using them to guide action generation. WING achieves average success rates of 99.20% on LIBERO, 93.80% on RoboTwin 2.0, and 57.7% on RoboCasa-GR1, and also performs strongly across four real-world manipulation tasks under diverse generalization settings. These results show that interaction-centric spectral guidance provides an effective and scalable way to transfer physical interaction knowledge from human egocentric video to robot control. Project page: https://mikuz12.github.io/wing/