🤖 AI Summary
Existing quadruped motion generation methods rely on specialized motion capture or complex control schemes, making it challenging to efficiently reproduce the nuanced details of natural locomotion. This work proposes a two-stage diffusion model that requires only quadruped motion data for training and enables end-to-end translation from generic human actions to realistic, controllable quadruped behaviors. By leveraging structured conditional guidance and a motion sequence inpainting strategy, the approach achieves unprecedented fine-grained control over the head and individual limbs. It significantly outperforms current motion retargeting methods in both realism and controllability, offering a practical solution for applications in animation and virtual production.
📝 Abstract
Realistic animal motion for virtual production is typically obtained either through motion capture of highly trained performers who accurately mimic animal behavior, or by retargeting ordinary human motion using complex control setups. Both approaches are challenging and often fail to fully reproduce the nuances of natural animal motion, motivating data-driven alternatives. We present an automatic human-to-quadruped puppeteering framework that produces plausible and controllable quadruped motions from ordinary human motion data. Our approach employs a two-stage generative diffusion model trained purely on quadruped motion data. By introducing a structured conditioning and inpainting strategy, our method supports a wide range of actions, including walking, running, jumping, sitting, and lying. Furthermore, we enable fine-grained intuitive control of the quadruped motion such as head movement control and individual limb puppeteering. Experimental results demonstrate improved motion realism and controllability compared to existing retargeting approaches, highlighting the effectiveness of our framework as a tool for animation and virtual production applications.