BeatDance: Generating Beat-Consistent 3D Dance with Hierarchical Spatial-Temporal Modeling

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of precisely aligning dance movements with musical rhythms by proposing a diffusion-based music-to-dance generation framework. The method introduces a novel hierarchical decoupled attention module that effectively disentangles pose features from temporal dynamics. Furthermore, it incorporates an auxiliary dance-to-music module alongside a cycle-consistency learning mechanism, which leverages reconstruction discrepancies to reinforce loss signals and significantly enhance audio-dance synchronization. Experimental evaluations on two benchmark datasets demonstrate that the proposed approach outperforms existing state-of-the-art methods, offering an effective solution for generating highly coherent music-driven dance sequences.
📝 Abstract
Generating realistic 3D dance from music is a challenging task that requires accurate synchronization with musical rhythms while capturing the spatial complexity of human motion. Although existing methods can generate physically plausible dance motions, they often struggle to achieve precise alignment with music, such as the beat. To address this limitation, we propose a novel diffusion-based framework, BeatDance, with two components: 1) We present a Hierarchical Decoupled Attention (HDA) module, which first disentangles the learning of human pose and temporal dynamics. A hierarchical structure is then employed to capture both short-term and long-term dependencies, thereby enhancing spatial-temporal modeling. 2) We adopt cycle-consistent learning by introducing an auxiliary dance-to-music module. During training, discrepancies between the reconstructed and original music induce a stronger loss signal, effectively encouraging the consistency property between the music and dance motion. Extensive experimental results demonstrate that our proposed approach outperforms recent competitive methods on two benchmark datasets.
Problem

Research questions and friction points this paper is trying to address.

3D dance generation
music-to-dance
beat synchronization
spatial-temporal modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diffusion Model
Hierarchical Decoupled Attention
Cycle-Consistent Learning
3D Dance Generation
Spatio-Temporal Modeling
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xiaojian Shen
College of Software, Jilin University, Changchun, 130012, China
Dahu Shi
Dahu Shi
Hikvision Research Insititute
Computer VisionArtificial Intelligence
J
Jianrong Zhang
ReLER, AAII, University of Technology Sydney, Sydney, Australia
H
Hai Li
College of Computer Science and Technology, Jilin University, Changchun, 130012, China
H
Hongwei Zhao
College of Computer Science and Technology, Jilin University, Changchun, 130012, China
Dawei Zhang
Dawei Zhang
Zhejiang Normal University
Computer VisionDeep LearningMulti-modal Fusion
Yunzhi Zhuge
Yunzhi Zhuge
Dalian University of Technology
Computer Vision
Zhiliang Wu
Zhiliang Wu
Research Scientist, Siemens Technology
Representation learningMachine learningGaussian ProcessesHealthcare
G
Guanghui Yue
School of Biomedical Engineering, Shenzhen University Medical School, Shenzhen University, Shenzhen, 518060, China
Wei Zhou
Wei Zhou
Huazhong University of Science and Technology
IoT SecuritySystem SecurityHardware Security