ASurvey: Spatiotemporal Consistency in Video Generation

📅 2025-02-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses the core challenge of spatiotemporal inconsistency in video generation—specifically, the lack of inter-frame motion coherence and spatial structural stability. We systematically survey technical approaches across five dimensions: foundational architectures, information representations, generative paradigms, post-processing techniques, and evaluation metrics—establishing, for the first time, a comprehensive taxonomy for spatiotemporal consistency in video generation. Our analysis reveals the intrinsic mechanisms by which diffusion models, Transformers, optical-flow guidance, temporal interpolation, and consistency regularization enable effective motion modeling and structural preservation. We unify multi-dimensional evaluation metrics—including TVD and FVD—into an extensible consistency assessment protocol. The study identifies key bottlenecks in current methods and outlines three critical future directions: controllable temporal modeling, implicit motion disentanglement, and standardized, unified evaluation benchmarks.

Technology Category

Computer Vision: Video Understanding & Activity AnalysisKnowledge Representation and Reasoning: Geometric, Spatial, and Temporal ReasoningMachine Learning: Deep Generative Models & Autoencoders

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsSocial Networks and Social Media: Generative AI / large language models and their impact on social systemsGraph Algorithms and Modeling for the Web: Algorithms and analysis for heterogeneous, signed, attributed, multi-relational, temporal, higher-order, and annotated Web-related graphs
📝 Abstract
Video generation, by leveraging a dynamic visual generation method, pushes the boundaries of Artificial Intelligence Generated Content (AIGC). Video generation presents unique challenges beyond static image generation, requiring both high-quality individual frames and temporal coherence to maintain consistency across the spatiotemporal sequence. Recent works have aimed at addressing the spatiotemporal consistency issue in video generation, while few literature review has been organized from this perspective. This gap hinders a deeper understanding of the underlying mechanisms for high-quality video generation. In this survey, we systematically review the recent advances in video generation, covering five key aspects: foundation models, information representations, generation schemes, post-processing techniques, and evaluation metrics. We particularly focus on their contributions to maintaining spatiotemporal consistency. Finally, we discuss the future directions and challenges in this field, hoping to inspire further efforts to advance the development of video generation.
Problem

Research questions and friction points this paper is trying to address.

Addressing spatiotemporal consistency in video generation
Reviewing advances in video generation techniques
Identifying gaps in high-quality video generation research
Innovation

Methods, ideas, or system contributions that make the work stand out.

Dynamic visual generation method
Spatiotemporal consistency maintenance
Systematic review of video generation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhiyu Yin
School of Computer Science and Technology, Harbin Institute of Technology, Shenzhen
Kehai Chen
Kehai Chen
Harbin Institute of Technolgy (Shenzhen)
LLMNatural Language ProcessingAgentMulti-model Generation
Xuefeng Bai
Xuefeng Bai
Harbin Institute of Technology (Shenzhen)
Natural language processingSemanticsDialogue
R
Ruili Jiang
School of Computer Science and Technology, Harbin Institute of Technology, Shenzhen
Juntao Li
Juntao Li
Soochow University
Language ModelsText Generation
H
Hongdong Li
School of Computer Science and Technology, Central South University
J
Jin Liu
School of Computer Science and Technology, Central South University
Y
Yang Xiang
Peng Cheng Laboratory
J
Jun Yu
School of Computer Science and Technology, Harbin Institute of Technology, Shenzhen
M
Min Zhang
School of Computer Science and Technology, Harbin Institute of Technology, Shenzhen