Exploring the Evolution of Physics Cognition in Video Generation: A Survey

📅 2025-03-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Current video generation models often produce physically implausible outputs—visually realistic yet violating fundamental physical laws. To address this, we propose the first “three-tier physical cognition” taxonomy—comprising basic schema perception, passive physical knowledge acquisition, and active world simulation—grounded in cognitive science to systematically characterize the evolution of physical modeling in video generation. Methodologically, our framework integrates diffusion-based architectures, physics-inspired motion representations, knowledge embedding mechanisms, and a multi-level physical consistency evaluation benchmark. Key contributions include: (1) establishing the first systematic review framework explicitly targeting physical cognition in video generation; (2) advancing the field from visual fidelity toward human-like physical understanding; and (3) clarifying the synergistic progression path among interpretability, controllability, and physical consistency—providing a structured roadmap for both academic research and industrial development.

Technology Category

Computer Vision: Low Level & Physics-based VisionCognitive Modeling & Cognitive Systems: Simulating Human BehaviorNatural Language Processing: Generation

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsUser Modeling, Personalization and Recommendation: User modeling and simulation for interactive and conversational systemsSocial Networks and Social Media: Generative AI / large language models and their impact on social systems
📝 Abstract
Recent advancements in video generation have witnessed significant progress, especially with the rapid advancement of diffusion models. Despite this, their deficiencies in physical cognition have gradually received widespread attention - generated content often violates the fundamental laws of physics, falling into the dilemma of ''visual realism but physical absurdity". Researchers began to increasingly recognize the importance of physical fidelity in video generation and attempted to integrate heuristic physical cognition such as motion representations and physical knowledge into generative systems to simulate real-world dynamic scenarios. Considering the lack of a systematic overview in this field, this survey aims to provide a comprehensive summary of architecture designs and their applications to fill this gap. Specifically, we discuss and organize the evolutionary process of physical cognition in video generation from a cognitive science perspective, while proposing a three-tier taxonomy: 1) basic schema perception for generation, 2) passive cognition of physical knowledge for generation, and 3) active cognition for world simulation, encompassing state-of-the-art methods, classical paradigms, and benchmarks. Subsequently, we emphasize the inherent key challenges in this domain and delineate potential pathways for future research, contributing to advancing the frontiers of discussion in both academia and industry. Through structured review and interdisciplinary analysis, this survey aims to provide directional guidance for developing interpretable, controllable, and physically consistent video generation paradigms, thereby propelling generative models from the stage of ''visual mimicry'' towards a new phase of ''human-like physical comprehension''.
Problem

Research questions and friction points this paper is trying to address.

Addressing physical absurdity in AI-generated videos
Integrating physics cognition into video generation systems
Developing physically consistent video generation paradigms
Innovation

Methods, ideas, or system contributions that make the work stand out.

Integrate motion representations into generative systems
Use physical knowledge for realistic video generation
Propose three-tier taxonomy for physical cognition
🔎 Similar Papers
M
Minghui Lin
School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan, China
X
Xiang Wang
School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan, China
Y
Yishan Wang
School of Engineering, Westlake University, Hangzhou, China
S
Shu Wang
School of Control Science and Engineering, Shandong University, Jinan, China
F
Fengqi Dai
School of Engineering, Westlake University, Hangzhou, China
Pengxiang Ding
Pengxiang Ding
Zhejiang University
Human Motion PredictionLarge Language ModelEmbodied AI
Cunxiang Wang
Cunxiang Wang
Tsinghua University; ZhipuAI
Large Language ModelsLLM EvaluationLLM Post-training
Z
Zhengrong Zuo
School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan, China
Nong Sang
Nong Sang
Huazhong University of Science and Technology
Computer Vision and Pattern Recognition
Siteng Huang
Siteng Huang
Alibaba DAMO Academy | ZJU | Westlake University
Vision-language ModelsGenerative ModelsEmbodied AI
D
Donglin Wang
School of Engineering, Westlake University, Hangzhou, China