latent diffusion video modeling

Designs and implements video generative pipelines that compress raw frames into compact latent codes using representation autoencoders and train diffusion models in that latent space to sample, predict, or interpolate future video latents. Engineers and evaluates model components, loss functions, and inference strategies to produce stable long‑horizon rollouts and meet runtime/throughput constraints (e.g., real‑time generation).

latentdiffusionvideomodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.37
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the instability and degraded generation performance in video variational autoencoders when used with latent diffusion models, which arises from an excessive number of latent channels. To mitigate this issue, the authors propose a frequency-aware latent space compression method that selectively attenuates high-frequency components in the video latent representations, replacing conventional channel pruning. This approach preserves essential structural information while achieving the same compression ratio. The proposed method substantially improves reconstruction fidelity, enhances training stability of the diffusion model, and outperforms strong baseline methods in generation quality, thereby demonstrating the efficacy and advantages of frequency-guided compression for video generation tasks.

generative performancelatent channelslatent diffusion models

Four-Plane Factorized Video Autoencoders

Dec 05, 2024
MS
M. Suhail
🏛️ Google | University of British Columbia | Vector Institute for AI

To address the inefficiency in modeling high-dimensional latent spaces and the substantial computational overhead during training and inference in video generation, this paper proposes the Four-Plane Variational Autoencoder (4P-VAE). The method introduces a novel four-plane factorized latent space architecture, projecting spatiotemporal video volumes onto four orthogonal 2D planes. This design enables sublinear growth of latent dimensionality with respect to input resolution while preserving representation fidelity under high compression ratios. 4P-VAE natively supports diverse downstream tasks—including class-conditional generation, frame prediction, and video interpolation—and integrates seamlessly with latent diffusion models (LDMs) for joint training. Experiments demonstrate that 4P-VAE achieves high-fidelity video reconstruction while significantly accelerating LDM training and inference and reducing GPU memory consumption. Overall, it establishes a new paradigm for efficient latent-space modeling of high-dimensional temporal data.

Challenges in training latent variable models for videosDesigning autoencoders for compressed yet rich video representationsEfficient generative modeling for high-dimensional video data

Compressing Scene Dynamics: A Generative Approach

Oct 13, 2024
SY
Shanzhi Yin
🏛️ City University of Hong Kong | DAMO Academy, Alibaba Group

This work addresses ultra-low-bitrate compression of dynamically varying scene videos. Instead of conventional content-based modeling, it leverages natural motion patterns—such as flower swaying or boat drifting—as priors. Methodologically, it introduces, for the first time, a lightweight generative prior for scene motion, establishing a novel framework comprising dense motion representation, sparse motion coding, and optical-flow-guided diffusion decoding—fully abandoning inter-frame prediction. Key contributions are: (1) learning compact, generalizable motion priors from common dynamic scenes; and (2) designing a flow-driven generative decoding mechanism enabling high-fidelity dynamic reconstruction. Experiments demonstrate substantial gains over VVC across diverse dynamic sequences, maintaining strong motion consistency and visual quality at ultra-low bitrates (0.01–0.1 bpp), with comprehensive improvements in rate-distortion performance.

Achieves superior rate-distortion performance over ECMCompresses scene dynamics using motion pattern priorsEnables ultra-low bitrate video communication

Commercial video generation models produce high-fidelity results but remain inaccessible due to prohibitive training and inference costs, especially for high-resolution video synthesis. Method: We propose an image-conditioned VAE that compresses video into an ultra-compact motion latent space—achieving 64× latent compression—and integrate it with a two-stage diffusion architecture (text → image → video) to efficiently generate 1024×1024 videos. Our approach introduces a novel motion-content disentangled latent representation and models temporal redundancy as sparse dynamic changes. Contribution/Results: This is the first method to generate 1024×1024 videos in just 15.5 seconds on a single A100 GPU. Training completes in only 3,200 GPU-hours. The framework achieves state-of-the-art visual quality while drastically reducing computational overhead, significantly enhancing efficiency and scalability for high-resolution video generation.

Enhancing efficiency in high-resolution video synthesisMinimizing training and inference time for video LDMsReducing video generation costs via compressed motion latents

Generative Inbetweening: Adapting Image-to-Video Models for Keyframe Interpolation

Aug 27, 2024
XW
Xiaojuan Wang
🏛️ University of Washington | Google | UC Berkeley

To address motion discontinuity and poor temporal consistency in keyframe-based video interpolation, this paper proposes a lightweight bidirectional diffusion sampling framework. Without retraining large-scale models, it fine-tunes pre-trained image-to-video diffusion models (e.g., Sora-like architectures) to enable bidirectional temporal modeling. The method initiates collaborative sampling from both end keyframes and introduces an overlapping estimation fusion strategy to enhance motion plausibility and structural fidelity of intermediate frames. To our knowledge, this is the first work to efficiently adapt unidirectional image-to-video diffusion models for keyframe interpolation. Extensive experiments demonstrate that our approach significantly outperforms optical-flow-based methods and existing diffusion-based interpolation techniques across multiple benchmarks, achieving state-of-the-art performance in visual quality, motion smoothness, and temporal consistency.

Adapting image-to-video modelsDual-directional diffusion sampling processGenerating video sequences between keyframes

Latest Papers

What's happening recently
View more

Existing video generation methods based on latent diffusion models suffer from high computational costs, hindering real-time applications. This work introduces, for the first time, the concept of inter-frame redundancy reduction—long employed in conventional video compression—into diffusion Transformer architectures. The authors propose an inter-frame latent pruning strategy that requires no additional training and design an attention restoration mechanism to mitigate visual artifacts introduced by pruning. This approach effectively bridges the inconsistency between training and inference, achieving a video editing throughput of 12.44 FPS on an NVIDIA RTX 6000 GPU—1.44× faster than the baseline—while preserving generation quality.

computational efficiencylatent diffusion modelsreal-time applications

Existing video variational autoencoders (VVAEs) struggle to simultaneously achieve high reconstruction fidelity, temporal coherence in latent representations, and diffusibility, thereby limiting generative performance. This work introduces predictive world modeling into the VVAE framework for the first time, proposing a unified reconstruction–prediction objective: encoding only a subset of past frames to jointly reconstruct observed frames and predict future ones. A stochastic future-frame dropout strategy is incorporated during training to end-to-end optimize the latent dynamics structure. The proposed approach substantially enhances both temporal consistency and generative capability of the latent representations, achieving a 34.42 improvement in FVD over Wan2.2 VAE on UCF101, accelerating convergence by 52%, and consistently yielding performance gains in downstream video understanding tasks.

diffusabilitylatent representationpredictive learning

This study addresses the reconstruction degradation and distribution shift issues caused by highly compressed video autoencoders when integrated with pretrained Diffusion Transformers (DiTs), which hinder efficient video generation. We introduce a novel generation-aware latent compression mechanism within a two-stage adaptation framework. Specifically, we freeze base latents while learning residual latents aligned in the DiT feature space to eliminate distribution shift, followed by lightweight fine-tuning combined with an asymmetric denoising strategy for efficient transfer. Crucially, our approach avoids training from scratch. Applied to Wan2.1-I2V-14B, it achieves an 8× token reduction and an 11.1× latency decrease while preserving VBench generation quality comparable to the original model.

Diffusion TransformerGeneration CompatibilityLatent Compression

Existing 3D variational autoencoders struggle to effectively model the semantic content and spatiotemporal structure of videos, and the rich representations from frozen video foundation models (VFMs) have not yet been efficiently transformed into compact, generation-friendly latent spaces. This work proposes VideoRAE, which demonstrates for the first time that multiscale features from a frozen VFM can be efficiently compressed—via a lightweight 1D self-attention projector—into either continuous or discrete latent representations. To enhance semantic fidelity during decoding, VideoRAE introduces a local-global alignment objective. Notably, the method operates without KL regularization and is compatible with both diffusion and autoregressive architectures. It achieves state-of-the-art class-conditional video generation performance on UCF-101 (gFVD: 40 for AR, 93 for DiT), converges approximately five times faster than existing autoencoders, and significantly accelerates training convergence in 2B-scale text-to-video tasks.

Generative ModelingLatent SpaceRepresentation Autoencoders

This work addresses the challenge of achieving both high perceptual quality and temporal consistency in video compression at extremely low bitrates (<0.005 bpp), where existing methods struggle. The authors propose a unified generative framework that introduces a causal tokenizer to decompose latent representations into I-latents and P-latents, and employs a Group-of-Latents strategy for structured modeling. Key latents are efficiently encoded via an I-frame Deep Compression Module (I-DCM), while a unified latent denoising module (U-LDM), built upon a pre-trained Diffusion Transformer (DiT), reconstructs high-fidelity intra-frame textures and coherent temporal dynamics directly from noise. Notably, this approach incurs no additional bitrate overhead and significantly outperforms current state-of-the-art methods under extreme low-bitrate constraints, delivering spatially detailed and temporally stable visual quality.

extreme low-bitratelatent generative modelingperceptual video compression

Hot Scholars

PW

Pengfei Wan

Head of Kling Video Generation Models, Kuaishou Technology
Generative ModelsComputer VisionMultimodal AIComputer Graphics
FJ

Felix Juefei-Xu

Research Scientist, Meta Superintelligence Labs
Generative ModelsDeep LearningComputer VisionAI Safety
JH

Ji Hou

Research Scientist, Meta Superintelligence Labs
Generative AI3D Computer Vision
VJ

Varun Jampani

Vice President of Research, Stability AI
Computer VisionMachine Learning