Score
Designs and implements video generative pipelines that compress raw frames into compact latent codes using representation autoencoders and train diffusion models in that latent space to sample, predict, or interpolate future video latents. Engineers and evaluates model components, loss functions, and inference strategies to produce stable long‑horizon rollouts and meet runtime/throughput constraints (e.g., real‑time generation).
This survey systematically addresses key challenges in diffusion-based video generation: temporal inconsistency, high computational cost, and ethical risks. We propose a fine-grained methodological taxonomy—first unifying temporal consistency modeling, efficient training strategies, and ethical governance within a single analytical framework. Dedicated sections comprehensively review evaluation metrics, industrial-grade deployment pipelines, and engineering best practices. We further integrate emerging directions—including video representation learning, motion modeling, video super-resolution, and cross-modal synergies (e.g., video question answering and retrieval)—to bridge theoretical advances with real-world implementation. Compared to existing surveys, ours offers broader coverage, greater novelty, and deeper technical insight. To support reproducibility and community advancement, we open-source a structured literature repository. This work serves as an authoritative reference and practical guide for researchers and engineers in generative video research and development.
This work addresses the instability and degraded generation performance in video variational autoencoders when used with latent diffusion models, which arises from an excessive number of latent channels. To mitigate this issue, the authors propose a frequency-aware latent space compression method that selectively attenuates high-frequency components in the video latent representations, replacing conventional channel pruning. This approach preserves essential structural information while achieving the same compression ratio. The proposed method substantially improves reconstruction fidelity, enhances training stability of the diffusion model, and outperforms strong baseline methods in generation quality, thereby demonstrating the efficacy and advantages of frequency-guided compression for video generation tasks.
To address the inefficiency in modeling high-dimensional latent spaces and the substantial computational overhead during training and inference in video generation, this paper proposes the Four-Plane Variational Autoencoder (4P-VAE). The method introduces a novel four-plane factorized latent space architecture, projecting spatiotemporal video volumes onto four orthogonal 2D planes. This design enables sublinear growth of latent dimensionality with respect to input resolution while preserving representation fidelity under high compression ratios. 4P-VAE natively supports diverse downstream tasks—including class-conditional generation, frame prediction, and video interpolation—and integrates seamlessly with latent diffusion models (LDMs) for joint training. Experiments demonstrate that 4P-VAE achieves high-fidelity video reconstruction while significantly accelerating LDM training and inference and reducing GPU memory consumption. Overall, it establishes a new paradigm for efficient latent-space modeling of high-dimensional temporal data.
This work addresses ultra-low-bitrate compression of dynamically varying scene videos. Instead of conventional content-based modeling, it leverages natural motion patterns—such as flower swaying or boat drifting—as priors. Methodologically, it introduces, for the first time, a lightweight generative prior for scene motion, establishing a novel framework comprising dense motion representation, sparse motion coding, and optical-flow-guided diffusion decoding—fully abandoning inter-frame prediction. Key contributions are: (1) learning compact, generalizable motion priors from common dynamic scenes; and (2) designing a flow-driven generative decoding mechanism enabling high-fidelity dynamic reconstruction. Experiments demonstrate substantial gains over VVC across diverse dynamic sequences, maintaining strong motion consistency and visual quality at ultra-low bitrates (0.01–0.1 bpp), with comprehensive improvements in rate-distortion performance.
Commercial video generation models produce high-fidelity results but remain inaccessible due to prohibitive training and inference costs, especially for high-resolution video synthesis. Method: We propose an image-conditioned VAE that compresses video into an ultra-compact motion latent space—achieving 64× latent compression—and integrate it with a two-stage diffusion architecture (text → image → video) to efficiently generate 1024×1024 videos. Our approach introduces a novel motion-content disentangled latent representation and models temporal redundancy as sparse dynamic changes. Contribution/Results: This is the first method to generate 1024×1024 videos in just 15.5 seconds on a single A100 GPU. Training completes in only 3,200 GPU-hours. The framework achieves state-of-the-art visual quality while drastically reducing computational overhead, significantly enhancing efficiency and scalability for high-resolution video generation.
To address motion discontinuity and poor temporal consistency in keyframe-based video interpolation, this paper proposes a lightweight bidirectional diffusion sampling framework. Without retraining large-scale models, it fine-tunes pre-trained image-to-video diffusion models (e.g., Sora-like architectures) to enable bidirectional temporal modeling. The method initiates collaborative sampling from both end keyframes and introduces an overlapping estimation fusion strategy to enhance motion plausibility and structural fidelity of intermediate frames. To our knowledge, this is the first work to efficiently adapt unidirectional image-to-video diffusion models for keyframe interpolation. Extensive experiments demonstrate that our approach significantly outperforms optical-flow-based methods and existing diffusion-based interpolation techniques across multiple benchmarks, achieving state-of-the-art performance in visual quality, motion smoothness, and temporal consistency.
Existing video generation methods based on latent diffusion models suffer from high computational costs, hindering real-time applications. This work introduces, for the first time, the concept of inter-frame redundancy reduction—long employed in conventional video compression—into diffusion Transformer architectures. The authors propose an inter-frame latent pruning strategy that requires no additional training and design an attention restoration mechanism to mitigate visual artifacts introduced by pruning. This approach effectively bridges the inconsistency between training and inference, achieving a video editing throughput of 12.44 FPS on an NVIDIA RTX 6000 GPU—1.44× faster than the baseline—while preserving generation quality.
Existing video variational autoencoders (VVAEs) struggle to simultaneously achieve high reconstruction fidelity, temporal coherence in latent representations, and diffusibility, thereby limiting generative performance. This work introduces predictive world modeling into the VVAE framework for the first time, proposing a unified reconstruction–prediction objective: encoding only a subset of past frames to jointly reconstruct observed frames and predict future ones. A stochastic future-frame dropout strategy is incorporated during training to end-to-end optimize the latent dynamics structure. The proposed approach substantially enhances both temporal consistency and generative capability of the latent representations, achieving a 34.42 improvement in FVD over Wan2.2 VAE on UCF101, accelerating convergence by 52%, and consistently yielding performance gains in downstream video understanding tasks.
This study addresses the reconstruction degradation and distribution shift issues caused by highly compressed video autoencoders when integrated with pretrained Diffusion Transformers (DiTs), which hinder efficient video generation. We introduce a novel generation-aware latent compression mechanism within a two-stage adaptation framework. Specifically, we freeze base latents while learning residual latents aligned in the DiT feature space to eliminate distribution shift, followed by lightweight fine-tuning combined with an asymmetric denoising strategy for efficient transfer. Crucially, our approach avoids training from scratch. Applied to Wan2.1-I2V-14B, it achieves an 8× token reduction and an 11.1× latency decrease while preserving VBench generation quality comparable to the original model.
Existing 3D variational autoencoders struggle to effectively model the semantic content and spatiotemporal structure of videos, and the rich representations from frozen video foundation models (VFMs) have not yet been efficiently transformed into compact, generation-friendly latent spaces. This work proposes VideoRAE, which demonstrates for the first time that multiscale features from a frozen VFM can be efficiently compressed—via a lightweight 1D self-attention projector—into either continuous or discrete latent representations. To enhance semantic fidelity during decoding, VideoRAE introduces a local-global alignment objective. Notably, the method operates without KL regularization and is compatible with both diffusion and autoregressive architectures. It achieves state-of-the-art class-conditional video generation performance on UCF-101 (gFVD: 40 for AR, 93 for DiT), converges approximately five times faster than existing autoencoders, and significantly accelerates training convergence in 2B-scale text-to-video tasks.
This work addresses the challenge of achieving both high perceptual quality and temporal consistency in video compression at extremely low bitrates (<0.005 bpp), where existing methods struggle. The authors propose a unified generative framework that introduces a causal tokenizer to decompose latent representations into I-latents and P-latents, and employs a Group-of-Latents strategy for structured modeling. Key latents are efficiently encoded via an I-frame Deep Compression Module (I-DCM), while a unified latent denoising module (U-LDM), built upon a pre-trained Diffusion Transformer (DiT), reconstructs high-fidelity intra-frame textures and coherent temporal dynamics directly from noise. Notably, this approach incurs no additional bitrate overhead and significantly outperforms current state-of-the-art methods under extreme low-bitrate constraints, delivering spatially detailed and temporally stable visual quality.