🤖 AI Summary
This study addresses the challenge of precisely aligning dance movements with musical rhythms by proposing a diffusion-based music-to-dance generation framework. The method introduces a novel hierarchical decoupled attention module that effectively disentangles pose features from temporal dynamics. Furthermore, it incorporates an auxiliary dance-to-music module alongside a cycle-consistency learning mechanism, which leverages reconstruction discrepancies to reinforce loss signals and significantly enhance audio-dance synchronization. Experimental evaluations on two benchmark datasets demonstrate that the proposed approach outperforms existing state-of-the-art methods, offering an effective solution for generating highly coherent music-driven dance sequences.
📝 Abstract
Generating realistic 3D dance from music is a challenging task that requires accurate synchronization with musical rhythms while capturing the spatial complexity of human motion. Although existing methods can generate physically plausible dance motions, they often struggle to achieve precise alignment with music, such as the beat. To address this limitation, we propose a novel diffusion-based framework, BeatDance, with two components: 1) We present a Hierarchical Decoupled Attention (HDA) module, which first disentangles the learning of human pose and temporal dynamics. A hierarchical structure is then employed to capture both short-term and long-term dependencies, thereby enhancing spatial-temporal modeling. 2) We adopt cycle-consistent learning by introducing an auxiliary dance-to-music module. During training, discrepancies between the reconstructed and original music induce a stronger loss signal, effectively encouraging the consistency property between the music and dance motion. Extensive experimental results demonstrate that our proposed approach outperforms recent competitive methods on two benchmark datasets.