🤖 AI Summary
This study addresses the limited tracking capability in humanoid robot interaction generation, which leads to low data utilization and physically implausible responses. To overcome these issues, we propose DIGHT, a framework that couples a digital interaction generator with a humanoid tracking policy. By leveraging simulated execution, it constructs physics-based preference pairs for joint optimization. Furthermore, this work introduces decoupled diffusion Direct Preference Optimization (DPO) to preserve multi-objective supervisory signals, and incorporates contact force feedback to quantify interaction fidelity while circumventing gradients through the simulator. Experimental results demonstrate that the proposed approach significantly enhances the physical plausibility of generated motions, achieving more reliable and faithful human-robot interactions.
📝 Abstract
Recent methods have made promising progress in generating interactions between two humanoids, largely relying on physics-based tracking policies to convert digital reference motions into executable trajectories. However, limited tracking capabilities restrict the range of reference motions that can be successfully executed, reducing data utilization. Moreover, even successful tracking does not guarantee physically plausible responses or faithful realization of the intended interactions. In this paper, we introduce DIGHT, a co-adaptive framework that couples a Digital human Interaction Generator with a Humanoid Tracking policy. Our DIGHT first executes multiple text-conditioned interaction candidates in simulation using a fixed tracker. It then constructs physics-grounded preferences from the resulting rollouts, covering both general executability and interaction fidelity. Rather than collapsing these signals into a single scalar reward for candidate ranking, we align the pretrained generator using physics-decoupled diffusion direct preference optimization (DPO), preserving criterion-specific supervision without differentiating through the simulator. To improve executability, preference pairs are derived from tracking error, friction, and floating. Additionally, to improve interaction fidelity, we propose to incorporate force feedback from simulator as a measure of contact fidelity and construct preferences over contact occurrence, location, duration, and force magnitude. The aligned generator then supplies reference motions for fine-tuning the tracker, improving compatibility between generation and physical execution. Extensive experiments demonstrate that our approach not only improves the physical plausibility of generated motions but also enables more reliable and faithful humanoid interactions in simulation.