🤖 AI Summary
This paper addresses the core challenge of spatiotemporal inconsistency in video generation—specifically, the lack of inter-frame motion coherence and spatial structural stability. We systematically survey technical approaches across five dimensions: foundational architectures, information representations, generative paradigms, post-processing techniques, and evaluation metrics—establishing, for the first time, a comprehensive taxonomy for spatiotemporal consistency in video generation. Our analysis reveals the intrinsic mechanisms by which diffusion models, Transformers, optical-flow guidance, temporal interpolation, and consistency regularization enable effective motion modeling and structural preservation. We unify multi-dimensional evaluation metrics—including TVD and FVD—into an extensible consistency assessment protocol. The study identifies key bottlenecks in current methods and outlines three critical future directions: controllable temporal modeling, implicit motion disentanglement, and standardized, unified evaluation benchmarks.
📝 Abstract
Video generation, by leveraging a dynamic visual generation method, pushes the boundaries of Artificial Intelligence Generated Content (AIGC). Video generation presents unique challenges beyond static image generation, requiring both high-quality individual frames and temporal coherence to maintain consistency across the spatiotemporal sequence. Recent works have aimed at addressing the spatiotemporal consistency issue in video generation, while few literature review has been organized from this perspective. This gap hinders a deeper understanding of the underlying mechanisms for high-quality video generation. In this survey, we systematically review the recent advances in video generation, covering five key aspects: foundation models, information representations, generation schemes, post-processing techniques, and evaluation metrics. We particularly focus on their contributions to maintaining spatiotemporal consistency. Finally, we discuss the future directions and challenges in this field, hoping to inspire further efforts to advance the development of video generation.