music-conditioned dance generation

Designs and builds models and pipelines that generate temporally coherent human dance or motion sequences and corresponding video conditioned on input music, mapping audio features to motion or visual outputs and aligning motion timing to beats and rhythm. Work includes selecting motion or visual representations (e.g., 3D skeletons, 2D keypoints, rendered frames), modeling style and genre diversity, controlling avatar or choreography constraints, and evaluating synchronization, realism, and diversity of generated outputs.

music-conditioneddancegeneration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.08
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

This study addresses the temporal-semantic inconsistencies arising from the isolated processing of audio and video in existing models by proposing a unified probabilistic framework that formulates joint generation, cross-modal generation, and editing as three subproblems under a single distribution. Building upon this framework, we construct a five-axis taxonomy that systematically organizes joint audio-visual editing tasks for the first time, mapping nine editing paradigms encompassing twenty-eight distinct types. By integrating multimodal fusion techniques with standardized evaluation protocols, this work identifies optimal methods, datasets, and metrics for various scenarios. Furthermore, it distills the most impactful open problems in the field, providing a comprehensive roadmap for future research in unified audio-visual editing.

audio-visual coherencecross-modal generationgenerative models

Must-Read Papers

Most classic and influential ideas
View more

MatchDance: Collaborative Mamba-Transformer Architecture Matching for High-Quality 3D Dance Synthesis

May 20, 2025
KY
Kaixing Yang
🏛️ Renmin University of China | Malou Tech Inc | Tsinghua University | Beihang University

Music-driven 3D dance generation suffers from insufficient choreographic consistency. To address this, we propose a two-stage collaborative framework: (1) a kinematic-dynamic-constrained Finite Scalar Quantization (FSQ) scheme that constructs high-fidelity, disentangled dance latent representations; and (2) a Mamba-Transformer hybrid architecture enabling precise cross-modal mapping from music to the learned latent space. We introduce the first kinematic-dynamic quantization paradigm synergized with Mamba-Transformer modeling, and establish the inaugural music–dance cross-modal retrieval evaluation framework tailored for dance generation, including dedicated metrics. Evaluated on the FineDance dataset, our method achieves state-of-the-art performance, significantly improving motion coherence, beat alignment, and 3D motion naturalness. The source code is publicly available.

Enhancing choreographic consistency in music-to-dance generationMapping music to latent dance representations accuratelyOvercoming limitations in 3D dance motion synthesis

This work addresses the limitations of existing dance video generation methods, which suffer from a scarcity of high-quality data and insufficient integration of music with generative models. To overcome these challenges, the authors introduce CIPE-Dance, a large-scale dataset comprising 300,000 high-fidelity dance video clips, and propose OmniDance—a unified framework capable of high-fidelity dance synthesis driven by text (TI2V), music (MI2V), or multimodal inputs (MTI2V). OmniDance innovatively integrates a depth-aware architecture, anchor-based curriculum learning, and a modality-specific time-varying classifier-free guidance (CFG) mechanism. Rigorous expert-guided data curation and annotation further enhance dataset quality. Evaluated on CIPE-Dance, OmniDance achieves state-of-the-art performance across all three generation tasks, significantly improving motion-music rhythmic alignment and visual fidelity.

dance video generationlarge-scale datasetmultimodal integration

This study addresses the challenge of simultaneously achieving motion alignment, identity preservation, and visual realism in music-driven dance generation. We propose a parallel pose-RGB dual-stream diffusion framework that integrates timestep-aware pose injection with a persistent identity mechanism to enable joint modeling of 3D motion and 2D visuals. Additionally, we construct a high-resolution in-the-wild dance dataset. By effectively combining explicit motion control with reference image synthesis, our method demonstrates superior performance in both dance generation and video synthesis tasks. The proposed approach significantly enhances temporal coherence and identity consistency, ultimately facilitating high-fidelity music-driven dance video generation.

Identity-preserving human animationMusic-driven dance video generationMusic-to-motion correspondence

X-Dancer: Expressive Music to Human Dance Video Generation

Feb 24, 2025
ZC
Zeyuan Chen
🏛️ UC San Diego | ByteDance | University of Southern California

This paper introduces the first zero-shot music-driven 2D dance video generation method, enabling long-duration, expressive, beat-aligned, and photorealistic dance videos from a single static portrait and arbitrary music. The approach employs a unified Transformer-diffusion framework: first, an autoregressive Transformer generates music-synchronized, tokenized 2D pose sequences using a spatially composable pose representation and a global attention mechanism that jointly encodes musical style and motion context; second, an AdaIN-conditioned diffusion model animates the pose sequence into photorealistic video frames. The entire pipeline is end-to-end differentiable and requires no fine-tuning or domain-specific training data. Quantitative and qualitative evaluations demonstrate state-of-the-art performance in motion diversity, expressiveness, and visual realism, with robust cross-style, long-sequence generation capability. The code and models are publicly released.

Enhances scalability using monocular video dataGenerates dance videos from static imagesSynchronizes 2D dance motions with music

Human Motion Video Generation: A Survey

Sep 04, 2025
HX
Haiwei Xue
🏛️ Tsinghua University | 01.AI | Xi’an Jiaotong University | University of Chinese Academy of Sciences | Nanjing University of Science and Technology | Shenzhen Institutes of Advanced Technology | Huawei Noah’s Ark Lab | Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ) | Artificial Intelligence Innovation and Incubation (AI) Institute of Fudan University | Shenzhen University | Carleton University

Existing surveys predominantly focus on isolated methodologies, lacking a systematic examination of the end-to-end pipeline for human motion video generation. To address this gap, we propose the first unified framework encompassing five core stages: input specification, motion planning, generation, optimization, and output rendering—supporting over ten subtasks driven by visual, textual, and audio modalities. Our work systematically reviews 200+ papers and constructs the field’s first comprehensive technical taxonomy. We innovatively investigate the potential of large language models (LLMs) for motion semantic modeling and cross-modal alignment. Furthermore, we integrate state-of-the-art techniques—including diffusion models, generative adversarial networks (GANs), and multimodal fusion—to identify key breakthroughs and release an open-source model library. This study fills a critical void in holistic, cross-cutting research on human motion video generation, providing both theoretical foundations and practical guidelines for applications such as digital avatars.

Addressing gaps across ten sub-tasks and five key phasesExploring large language models' potential for motion generationLack of comprehensive overview in human motion video generation

Latest Papers

What's happening recently
View more

Existing approaches struggle to effectively evaluate rhythmic coupling and cross-modal alignment in music-and-dance generation. This work proposes a multi-level evaluation framework that systematically assesses text-driven joint generation systems along three dimensions: single-modality quality, adherence to textual instructions, and music–dance rhythmic alignment. Innovatively integrating physically computable metrics with human perceptual judgments, the study introduces the first rhythm alignment dataset and a structured music semantic descriptor, alongside a unified baseline model named RhyJAM. Experiments reveal that prevailing audio-visual models exhibit notable deficiencies in beat synchronization, whereas RhyJAM significantly enhances beat-level cross-modal alignment while maintaining high-quality unimodal outputs.

audio-visual generationcross-modal synchronizationevaluation benchmark

Existing music-driven dance generation methods often overlook the compositional structure of movement, resulting in outputs that lack structural coherence and controllability. This work proposes modeling dance as a sequence of semantically interpretable atomic actions and introduces a two-stage generation framework: first planning the type, duration, and timing of atomic actions based on input music, then synthesizing them into smooth, coherent full-body motion. A reusable and editable motion vocabulary is constructed through large-scale motion segmentation, clustering, and semantic relabeling via large language models, enabling structure-aware dance synthesis. Experiments demonstrate that the proposed approach significantly outperforms existing methods in structural coherence, rhythmic alignment, and perceptual naturalness, while supporting flexible editing through its explicit structural representation.

atomic movementsdance controllabilitymotion compositionality

Existing 3D dance generation methods struggle to achieve fine-grained control over multimodal inputs such as music and text, resulting in limited expressiveness and misalignment with creative intent. This work proposes a coarse-to-fine interactive generation system inspired by professional choreographic workflows: it first leverages a multimodal large language model to interpret user prompts and retrieve high-quality motion clips, then employs a music-conditioned diffusion refiner to seamlessly connect and iteratively optimize the motion sequence. Introducing a human-centered interactive AI choreography paradigm, the approach incrementally integrates user intent throughout the generation pipeline, jointly enhancing controllability, expressiveness, and output quality. Experimental results demonstrate that the proposed system significantly outperforms existing methods in both quantitative and qualitative evaluations, effectively empowering users’ creative expression and practical choreographic utility.

3D dance generationAI-assisted choreographyexpressive motion

Existing methods struggle to generate minute-long, high-resolution dance videos synchronized with music, often hindered by the temporal limitations of diffusion models, which lead to temporal drift, identity inconsistency, and repetitive motions. This work proposes a hierarchical generation framework that decouples music-to-dance synthesis into global keyframe planning and local temporal refinement, leveraging full-song audio context to ensure long-term coherence while supporting dual conditioning on both audio and text. The approach introduces a novel time-mapping RoPE embedding with dynamic frame-rate adaptation for precise audio-motion alignment, incorporates an optical flow loss to enhance motion continuity, and integrates motion velocity control to preserve fine details of fast movements. To our knowledge, this is the first method capable of stably generating high-fidelity dance videos exceeding one minute in duration at 720p resolution and 30 fps, achieving state-of-the-art performance across five distinct dance styles.

long-duration video synthesismotion consistencymusic-to-dance generation

This work addresses the challenges of text-driven controllable dance generation, which are primarily hindered by the scarcity of high-quality data and the inherent complexity of dance motion—particularly its spatial dynamics, strong directional constraints, and highly decoupled movements across body parts. To overcome these limitations, the authors propose a theoretical framework termed “choreographic grammar,” integrating principles from dance theory, human anatomy, and biomechanics. They introduce DanceFlow, a novel dataset comprising 41 hours of high-fidelity motion capture paired with 6.34 million words of fine-grained textual descriptions, and develop DanceCrafter, a motion Transformer built upon the Momentum Human Rig skeleton. The model incorporates continuous manifold-based motion representations, hybrid normalization, and an anatomy-aware loss function. Quantitative evaluations and user studies demonstrate that this approach significantly outperforms existing methods in motion quality, fine-grained controllability, and naturalness of generation.

choreographic representationcomplex choreography modelingcontrollable motion synthesis

Hot Scholars

JP

Joseph P. Near

University of Vermont
Security & PrivacyDifferential PrivacyProgramming LanguagesFormal Methods
SX

Stella X. Yu

Professor of EECS, University of Michigan
Computer VisionRoboticsEmbodied AIDevelopmental AI
ZW

Zilin Wang

University of Oxford
Deep Reinforcement LearningAutonomous Driving
RS

Rodolphe Sepulchre

Professor of Engineering, KU Leuven and University of Cambridge
controloptimizationneuronal behaviors
CS

Christian Skalka

Associate Professor of Computer Science, University of Vermont
Programming LanguagesLogic in Computer ScienceWireless Embedded Systems