Score
Designs and builds models and pipelines that generate temporally coherent human dance or motion sequences and corresponding video conditioned on input music, mapping audio features to motion or visual outputs and aligning motion timing to beats and rhythm. Work includes selecting motion or visual representations (e.g., 3D skeletons, 2D keypoints, rendered frames), modeling style and genre diversity, controlling avatar or choreography constraints, and evaluating synchronization, realism, and diversity of generated outputs.
The field of motion generation lacks a systematic survey grounded in generative methodology. To address this gap, we propose the first deep taxonomy centered on generative strategies, synthesizing state-of-the-art works from top-tier conferences (CVPR, ICCV, SIGGRAPH, CoRL) since 2023. Our framework uniformly analyzes four dominant paradigms—GANs, autoencoders, autoregressive models, and diffusion models—across three dimensions: architectural design, conditional modeling mechanisms, and evaluation protocols for motion sequence synthesis. We consolidate widely adopted datasets and metrics, identifying key challenges including motion coherence, physical plausibility, and cross-domain generalization. By establishing a comprehensive, comparable analytical benchmark with well-defined dimensions, our survey significantly enhances methodological comparability and facilitates precise problem diagnosis. This work serves as a foundational reference for researchers advancing generative motion modeling.
This study addresses the temporal-semantic inconsistencies arising from the isolated processing of audio and video in existing models by proposing a unified probabilistic framework that formulates joint generation, cross-modal generation, and editing as three subproblems under a single distribution. Building upon this framework, we construct a five-axis taxonomy that systematically organizes joint audio-visual editing tasks for the first time, mapping nine editing paradigms encompassing twenty-eight distinct types. By integrating multimodal fusion techniques with standardized evaluation protocols, this work identifies optimal methods, datasets, and metrics for various scenarios. Furthermore, it distills the most impactful open problems in the field, providing a comprehensive roadmap for future research in unified audio-visual editing.
Music-driven 3D dance generation suffers from insufficient choreographic consistency. To address this, we propose a two-stage collaborative framework: (1) a kinematic-dynamic-constrained Finite Scalar Quantization (FSQ) scheme that constructs high-fidelity, disentangled dance latent representations; and (2) a Mamba-Transformer hybrid architecture enabling precise cross-modal mapping from music to the learned latent space. We introduce the first kinematic-dynamic quantization paradigm synergized with Mamba-Transformer modeling, and establish the inaugural music–dance cross-modal retrieval evaluation framework tailored for dance generation, including dedicated metrics. Evaluated on the FineDance dataset, our method achieves state-of-the-art performance, significantly improving motion coherence, beat alignment, and 3D motion naturalness. The source code is publicly available.
This work addresses the limitations of existing dance video generation methods, which suffer from a scarcity of high-quality data and insufficient integration of music with generative models. To overcome these challenges, the authors introduce CIPE-Dance, a large-scale dataset comprising 300,000 high-fidelity dance video clips, and propose OmniDance—a unified framework capable of high-fidelity dance synthesis driven by text (TI2V), music (MI2V), or multimodal inputs (MTI2V). OmniDance innovatively integrates a depth-aware architecture, anchor-based curriculum learning, and a modality-specific time-varying classifier-free guidance (CFG) mechanism. Rigorous expert-guided data curation and annotation further enhance dataset quality. Evaluated on CIPE-Dance, OmniDance achieves state-of-the-art performance across all three generation tasks, significantly improving motion-music rhythmic alignment and visual fidelity.
This study addresses the challenge of simultaneously achieving motion alignment, identity preservation, and visual realism in music-driven dance generation. We propose a parallel pose-RGB dual-stream diffusion framework that integrates timestep-aware pose injection with a persistent identity mechanism to enable joint modeling of 3D motion and 2D visuals. Additionally, we construct a high-resolution in-the-wild dance dataset. By effectively combining explicit motion control with reference image synthesis, our method demonstrates superior performance in both dance generation and video synthesis tasks. The proposed approach significantly enhances temporal coherence and identity consistency, ultimately facilitating high-fidelity music-driven dance video generation.
This paper introduces the first zero-shot music-driven 2D dance video generation method, enabling long-duration, expressive, beat-aligned, and photorealistic dance videos from a single static portrait and arbitrary music. The approach employs a unified Transformer-diffusion framework: first, an autoregressive Transformer generates music-synchronized, tokenized 2D pose sequences using a spatially composable pose representation and a global attention mechanism that jointly encodes musical style and motion context; second, an AdaIN-conditioned diffusion model animates the pose sequence into photorealistic video frames. The entire pipeline is end-to-end differentiable and requires no fine-tuning or domain-specific training data. Quantitative and qualitative evaluations demonstrate state-of-the-art performance in motion diversity, expressiveness, and visual realism, with robust cross-style, long-sequence generation capability. The code and models are publicly released.
Existing surveys predominantly focus on isolated methodologies, lacking a systematic examination of the end-to-end pipeline for human motion video generation. To address this gap, we propose the first unified framework encompassing five core stages: input specification, motion planning, generation, optimization, and output rendering—supporting over ten subtasks driven by visual, textual, and audio modalities. Our work systematically reviews 200+ papers and constructs the field’s first comprehensive technical taxonomy. We innovatively investigate the potential of large language models (LLMs) for motion semantic modeling and cross-modal alignment. Furthermore, we integrate state-of-the-art techniques—including diffusion models, generative adversarial networks (GANs), and multimodal fusion—to identify key breakthroughs and release an open-source model library. This study fills a critical void in holistic, cross-cutting research on human motion video generation, providing both theoretical foundations and practical guidelines for applications such as digital avatars.
Existing approaches struggle to effectively evaluate rhythmic coupling and cross-modal alignment in music-and-dance generation. This work proposes a multi-level evaluation framework that systematically assesses text-driven joint generation systems along three dimensions: single-modality quality, adherence to textual instructions, and music–dance rhythmic alignment. Innovatively integrating physically computable metrics with human perceptual judgments, the study introduces the first rhythm alignment dataset and a structured music semantic descriptor, alongside a unified baseline model named RhyJAM. Experiments reveal that prevailing audio-visual models exhibit notable deficiencies in beat synchronization, whereas RhyJAM significantly enhances beat-level cross-modal alignment while maintaining high-quality unimodal outputs.
Existing music-driven dance generation methods often overlook the compositional structure of movement, resulting in outputs that lack structural coherence and controllability. This work proposes modeling dance as a sequence of semantically interpretable atomic actions and introduces a two-stage generation framework: first planning the type, duration, and timing of atomic actions based on input music, then synthesizing them into smooth, coherent full-body motion. A reusable and editable motion vocabulary is constructed through large-scale motion segmentation, clustering, and semantic relabeling via large language models, enabling structure-aware dance synthesis. Experiments demonstrate that the proposed approach significantly outperforms existing methods in structural coherence, rhythmic alignment, and perceptual naturalness, while supporting flexible editing through its explicit structural representation.
Existing 3D dance generation methods struggle to achieve fine-grained control over multimodal inputs such as music and text, resulting in limited expressiveness and misalignment with creative intent. This work proposes a coarse-to-fine interactive generation system inspired by professional choreographic workflows: it first leverages a multimodal large language model to interpret user prompts and retrieve high-quality motion clips, then employs a music-conditioned diffusion refiner to seamlessly connect and iteratively optimize the motion sequence. Introducing a human-centered interactive AI choreography paradigm, the approach incrementally integrates user intent throughout the generation pipeline, jointly enhancing controllability, expressiveness, and output quality. Experimental results demonstrate that the proposed system significantly outperforms existing methods in both quantitative and qualitative evaluations, effectively empowering users’ creative expression and practical choreographic utility.
Existing methods struggle to generate minute-long, high-resolution dance videos synchronized with music, often hindered by the temporal limitations of diffusion models, which lead to temporal drift, identity inconsistency, and repetitive motions. This work proposes a hierarchical generation framework that decouples music-to-dance synthesis into global keyframe planning and local temporal refinement, leveraging full-song audio context to ensure long-term coherence while supporting dual conditioning on both audio and text. The approach introduces a novel time-mapping RoPE embedding with dynamic frame-rate adaptation for precise audio-motion alignment, incorporates an optical flow loss to enhance motion continuity, and integrates motion velocity control to preserve fine details of fast movements. To our knowledge, this is the first method capable of stably generating high-fidelity dance videos exceeding one minute in duration at 720p resolution and 30 fps, achieving state-of-the-art performance across five distinct dance styles.
This work addresses the challenges of text-driven controllable dance generation, which are primarily hindered by the scarcity of high-quality data and the inherent complexity of dance motion—particularly its spatial dynamics, strong directional constraints, and highly decoupled movements across body parts. To overcome these limitations, the authors propose a theoretical framework termed “choreographic grammar,” integrating principles from dance theory, human anatomy, and biomechanics. They introduce DanceFlow, a novel dataset comprising 41 hours of high-fidelity motion capture paired with 6.34 million words of fine-grained textual descriptions, and develop DanceCrafter, a motion Transformer built upon the Momentum Human Rig skeleton. The model incorporates continuous manifold-based motion representations, hybrid normalization, and an anatomy-aware loss function. Quantitative evaluations and user studies demonstrate that this approach significantly outperforms existing methods in motion quality, fine-grained controllability, and naturalness of generation.