residual motion latent diffusion

Designs and implements latent-diffusion generative models and pipelines that produce or edit temporally coherent motion sequences by modeling residual motion in a compact latent space and using cascaded diffusion stages to refine appearance and dynamics. This competence covers building architectures and training/evaluation procedures to estimate per-frame residual latent displacements, enforce strict temporal coherence, and synthesize or control dynamic sequences.

residualmotionlatentdiffusion

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.4
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

LaMD: Latent Motion Diffusion for Image-Conditional Video Generation

Apr 23, 2023
YH
Yaosi Hu
🏛️ Wuhan University | Microsoft Research Asia

Existing video diffusion models improve visual quality but suffer from poor motion coherence and low sampling efficiency. To address these limitations, we propose a two-stage image-conditioned video generation framework. First, a novel motion-decomposition video autoencoder disentangles implicit motion representations from appearance reconstruction. Second, a continuous latent-space diffusion model captures the image-conditioned motion prior. This approach enables expressive, multimodal, and computationally efficient motion modeling. Evaluated on BAIR, Landscape, NATOPS, MUG, and CATER-GEN benchmarks, our method significantly enhances motion naturalness and temporal consistency while accelerating sampling by 2.1–3.8×. Moreover, it supports complex dynamic modeling and fine-grained motion control.

Decompose video generation into latent motion modelingGenerate coherent motion in image-conditional videosImprove motion expressiveness and sampling efficiency

This work addresses the challenge of generating high-fidelity, natural, and diverse motion sequences from extremely sparse keyframes—a scenario where existing methods struggle to simultaneously ensure accuracy, temporal continuity, and variation. The authors propose a novel framework that integrates Implicit Neural Representations (INRs) with Latent Diffusion Models (LDMs), introducing continuous implicit representations into the diffusion generation paradigm for the first time. By sampling INR parameters under keyframe constraints, the method reconstructs plausible intermediate motions directly from minimal input. This approach significantly enhances generation quality in sparse keyframe settings, faithfully adhering to the given keyframes while preserving smoothness and semantic coherence throughout the synthesized motion sequence.

generative modelskeyframe preservationlatent diffusion

Generative Inbetweening: Adapting Image-to-Video Models for Keyframe Interpolation

Aug 27, 2024
XW
Xiaojuan Wang
🏛️ University of Washington | Google | UC Berkeley

To address motion discontinuity and poor temporal consistency in keyframe-based video interpolation, this paper proposes a lightweight bidirectional diffusion sampling framework. Without retraining large-scale models, it fine-tunes pre-trained image-to-video diffusion models (e.g., Sora-like architectures) to enable bidirectional temporal modeling. The method initiates collaborative sampling from both end keyframes and introduces an overlapping estimation fusion strategy to enhance motion plausibility and structural fidelity of intermediate frames. To our knowledge, this is the first work to efficiently adapt unidirectional image-to-video diffusion models for keyframe interpolation. Extensive experiments demonstrate that our approach significantly outperforms optical-flow-based methods and existing diffusion-based interpolation techniques across multiple benchmarks, achieving state-of-the-art performance in visual quality, motion smoothness, and temporal consistency.

Adapting image-to-video modelsDual-directional diffusion sampling processGenerating video sequences between keyframes

MotionLCM: Real-time Controllable Motion Generation via Latent Consistency Model

Apr 30, 2024
WD
Wen-Dao Dai
🏛️ Tsinghua University | Shanghai AI Laboratory

To address the slow inference speed and poor real-time applicability of existing spatiotemporal-controllable human motion generation methods, this paper proposes the first Motion Latent Consistency Model (MLCM) tailored for motion generation, coupled with a latent-space Motion ControlNet that incorporates explicit motion control signals for supervised training. Built upon a single-step or few-step sampling strategy within a motion latent diffusion framework, our approach enables highly efficient generation. Experiments demonstrate that MLCM achieves high-fidelity motion synthesis while maintaining precise joint control over textual prompts and initial motion conditions. Crucially, it attains an inference latency of <100 ms per frame—significantly outperforming existing diffusion-based methods—and marks the first realization of real-time, dual-driven (text + initial motion) controllable motion generation.

EfficiencyReal-time ControlVirtual Character Animation

MoLA: Motion Generation and Editing with Latent Diffusion Enhanced by Adversarial Training

Jun 04, 2024
KU
Kengo Uchida
🏛️ Sony AI | Sony Group Corporation

To address weak controllability, low generation quality, slow inference, and variable-length alignment challenges in text-to-motion synthesis, this paper proposes a unified framework. First, learnable activation variables are introduced to enable text-length-adaptive motion sequence generation. Second, an adversarially enhanced latent diffusion model (LDM) is constructed, incorporating Wasserstein adversarial training to improve motion realism. Third, a training-free classifier-free guidance mechanism is designed to support diverse motion editing—including start/end positions and pelvis trajectories. Built upon a joint VAE-LDM architecture, the method enables versatile control without additional fine-tuning. Experiments demonstrate significant improvements: 21.3% reduction in FID (indicating higher fidelity) and 3.2× faster inference speed. Notably, this is the first single-model approach to simultaneously achieve variable-length alignment, strong-constraint editing, and high-fidelity motion synthesis.

Enhances text-to-motion generation controllability.Facilitates multiple motion editing tasks efficiently.Improves motion generation quality and speed.

Latest Papers

What's happening recently
View more

This work addresses the limited controllability of temporal dynamics and editing in existing video diffusion Transformer models. To overcome this, the authors propose a lightweight temporal control module that enables explicit and fine-grained manipulation of motion speed and temporal structure without altering the pre-trained DiT backbone. By effectively leveraging the generative priors learned during pre-training, the method significantly enhances temporal controllability in video generation while preserving the original output quality. The approach thus offers a practical and efficient solution for precise temporal editing in diffusion-based video synthesis.

diffusion transformersmotion dynamicstemporal control

This study systematically compares diffusion models and rectified flows in the context of text-driven motion generation within a continuous latent space. Leveraging a unified MotionGPT3 architecture and consistent training protocols, we conduct controlled experiments on the HumanML3D dataset and demonstrate, for the first time, the advantages of rectified flows for this task. Specifically, rectified flows exhibit faster training convergence and superior early-stage generation quality, achieving motion fidelity on par with or exceeding that of diffusion models under identical conditions. Notably, they maintain stable and competitive performance even with significantly fewer sampling steps. These findings highlight the greater potential of rectified flows in balancing generation quality and computational efficiency for text-to-motion synthesis.

continuous latent spacediffusion modelsmotion priors

Existing video diffusion models struggle with temporally coherent and high-fidelity motion synthesis in highly dynamic scenes due to the limitations of static loss functions, which fail to adequately capture complex motion dynamics. To address this, this work proposes a latent temporal difference (LTD)-driven, motion-aware loss weighting strategy that leverages inter-frame changes in latent space as a motion prior. By assigning stronger penalties to regions exhibiting high temporal variation, the method stabilizes training and enhances the model’s capacity to reconstruct high-frequency dynamics. This approach overcomes the constraints of conventional static losses and achieves state-of-the-art performance, surpassing strong baselines by 3.31% on VBench and 3.58% on VMBench, thereby significantly improving motion fidelity in generated videos.

diffusion modelsdynamic fidelitymotion quality

This study investigates whether video diffusion models implicitly encode physical structure rather than merely reproducing motion patterns observed in training data. Through deterministic sampling inversion and latent trajectory reconstruction, combined with linear probing and attention analysis, the authors demonstrate that signals of physical plausibility emerge spontaneously within the denoising Transformer—not from the VAE input—and do so without any explicit self-supervised objective. The proposed approach achieves an average classification accuracy of 81.27% on the IntPhys and InfLevel benchmarks, substantially outperforming specialized representation learning baselines such as V-JEPA and VideoMAE.

latent trajectoriesphysical plausibilityrepresentation learning

Hot Scholars

LJ

Lutao Jiang

PhD Student at HKUST(GZ)
3D VisionComputer VisionGen AIAIGC
YC

Ying-Cong Chen

Hong Kong University of Science and Technology (Guangzhou)
Computer Vision and Pattern Recognition