Score
Designs and implements latent-diffusion generative models and pipelines that produce or edit temporally coherent motion sequences by modeling residual motion in a compact latent space and using cascaded diffusion stages to refine appearance and dynamics. This competence covers building architectures and training/evaluation procedures to estimate per-frame residual latent displacements, enforce strict temporal coherence, and synthesize or control dynamic sequences.
This survey systematically addresses key challenges in diffusion-based video generation: temporal inconsistency, high computational cost, and ethical risks. We propose a fine-grained methodological taxonomy—first unifying temporal consistency modeling, efficient training strategies, and ethical governance within a single analytical framework. Dedicated sections comprehensively review evaluation metrics, industrial-grade deployment pipelines, and engineering best practices. We further integrate emerging directions—including video representation learning, motion modeling, video super-resolution, and cross-modal synergies (e.g., video question answering and retrieval)—to bridge theoretical advances with real-world implementation. Compared to existing surveys, ours offers broader coverage, greater novelty, and deeper technical insight. To support reproducibility and community advancement, we open-source a structured literature repository. This work serves as an authoritative reference and practical guide for researchers and engineers in generative video research and development.
Existing video diffusion models improve visual quality but suffer from poor motion coherence and low sampling efficiency. To address these limitations, we propose a two-stage image-conditioned video generation framework. First, a novel motion-decomposition video autoencoder disentangles implicit motion representations from appearance reconstruction. Second, a continuous latent-space diffusion model captures the image-conditioned motion prior. This approach enables expressive, multimodal, and computationally efficient motion modeling. Evaluated on BAIR, Landscape, NATOPS, MUG, and CATER-GEN benchmarks, our method significantly enhances motion naturalness and temporal consistency while accelerating sampling by 2.1–3.8×. Moreover, it supports complex dynamic modeling and fine-grained motion control.
This work addresses the challenge of generating high-fidelity, natural, and diverse motion sequences from extremely sparse keyframes—a scenario where existing methods struggle to simultaneously ensure accuracy, temporal continuity, and variation. The authors propose a novel framework that integrates Implicit Neural Representations (INRs) with Latent Diffusion Models (LDMs), introducing continuous implicit representations into the diffusion generation paradigm for the first time. By sampling INR parameters under keyframe constraints, the method reconstructs plausible intermediate motions directly from minimal input. This approach significantly enhances generation quality in sparse keyframe settings, faithfully adhering to the given keyframes while preserving smoothness and semantic coherence throughout the synthesized motion sequence.
To address motion discontinuity and poor temporal consistency in keyframe-based video interpolation, this paper proposes a lightweight bidirectional diffusion sampling framework. Without retraining large-scale models, it fine-tunes pre-trained image-to-video diffusion models (e.g., Sora-like architectures) to enable bidirectional temporal modeling. The method initiates collaborative sampling from both end keyframes and introduces an overlapping estimation fusion strategy to enhance motion plausibility and structural fidelity of intermediate frames. To our knowledge, this is the first work to efficiently adapt unidirectional image-to-video diffusion models for keyframe interpolation. Extensive experiments demonstrate that our approach significantly outperforms optical-flow-based methods and existing diffusion-based interpolation techniques across multiple benchmarks, achieving state-of-the-art performance in visual quality, motion smoothness, and temporal consistency.
To address the slow inference speed and poor real-time applicability of existing spatiotemporal-controllable human motion generation methods, this paper proposes the first Motion Latent Consistency Model (MLCM) tailored for motion generation, coupled with a latent-space Motion ControlNet that incorporates explicit motion control signals for supervised training. Built upon a single-step or few-step sampling strategy within a motion latent diffusion framework, our approach enables highly efficient generation. Experiments demonstrate that MLCM achieves high-fidelity motion synthesis while maintaining precise joint control over textual prompts and initial motion conditions. Crucially, it attains an inference latency of <100 ms per frame—significantly outperforming existing diffusion-based methods—and marks the first realization of real-time, dual-driven (text + initial motion) controllable motion generation.
To address weak controllability, low generation quality, slow inference, and variable-length alignment challenges in text-to-motion synthesis, this paper proposes a unified framework. First, learnable activation variables are introduced to enable text-length-adaptive motion sequence generation. Second, an adversarially enhanced latent diffusion model (LDM) is constructed, incorporating Wasserstein adversarial training to improve motion realism. Third, a training-free classifier-free guidance mechanism is designed to support diverse motion editing—including start/end positions and pelvis trajectories. Built upon a joint VAE-LDM architecture, the method enables versatile control without additional fine-tuning. Experiments demonstrate significant improvements: 21.3% reduction in FID (indicating higher fidelity) and 3.2× faster inference speed. Notably, this is the first single-model approach to simultaneously achieve variable-length alignment, strong-constraint editing, and high-fidelity motion synthesis.
This work addresses the limited controllability of temporal dynamics and editing in existing video diffusion Transformer models. To overcome this, the authors propose a lightweight temporal control module that enables explicit and fine-grained manipulation of motion speed and temporal structure without altering the pre-trained DiT backbone. By effectively leveraging the generative priors learned during pre-training, the method significantly enhances temporal controllability in video generation while preserving the original output quality. The approach thus offers a practical and efficient solution for precise temporal editing in diffusion-based video synthesis.
This study systematically compares diffusion models and rectified flows in the context of text-driven motion generation within a continuous latent space. Leveraging a unified MotionGPT3 architecture and consistent training protocols, we conduct controlled experiments on the HumanML3D dataset and demonstrate, for the first time, the advantages of rectified flows for this task. Specifically, rectified flows exhibit faster training convergence and superior early-stage generation quality, achieving motion fidelity on par with or exceeding that of diffusion models under identical conditions. Notably, they maintain stable and competitive performance even with significantly fewer sampling steps. These findings highlight the greater potential of rectified flows in balancing generation quality and computational efficiency for text-to-motion synthesis.
Existing video diffusion models struggle with temporally coherent and high-fidelity motion synthesis in highly dynamic scenes due to the limitations of static loss functions, which fail to adequately capture complex motion dynamics. To address this, this work proposes a latent temporal difference (LTD)-driven, motion-aware loss weighting strategy that leverages inter-frame changes in latent space as a motion prior. By assigning stronger penalties to regions exhibiting high temporal variation, the method stabilizes training and enhances the model’s capacity to reconstruct high-frequency dynamics. This approach overcomes the constraints of conventional static losses and achieves state-of-the-art performance, surpassing strong baselines by 3.31% on VBench and 3.58% on VMBench, thereby significantly improving motion fidelity in generated videos.
This study investigates whether video diffusion models implicitly encode physical structure rather than merely reproducing motion patterns observed in training data. Through deterministic sampling inversion and latent trajectory reconstruction, combined with linear probing and attention analysis, the authors demonstrate that signals of physical plausibility emerge spontaneously within the denoising Transformer—not from the VAE input—and do so without any explicit self-supervised objective. The proposed approach achieves an average classification accuracy of 81.27% on the IntPhys and InfLevel benchmarks, substantially outperforming specialized representation learning baselines such as V-JEPA and VideoMAE.