PastNet: Introducing Physical Inductive Biases for Spatio-temporal Video Prediction

📅 2023-05-19
🏛️ ACM Multimedia
📈 Citations: 17
✨ Influential: 1
📄 PDF
🤖 AI Summary
This work addresses spatiotemporal prediction for high-resolution videos, aiming to efficiently and accurately synthesize future frames solely from historical frames and timestamps. We propose a physics-driven spectral-domain modeling framework: (1) physical priors are explicitly embedded into Fourier-frequency-domain convolutions to enforce physics-consistent inductive bias; and (2) a memory bank based on local intrinsic dimension estimation is introduced to enable feature discretization and compact storage. The resulting method is both lightweight and highly generalizable. Extensive evaluations on major video prediction benchmarks—including KTH, BAIR, and Moving MNIST—demonstrate substantial improvements over state-of-the-art methods: PSNR and SSIM increase by 1.2–2.8 dB, inference speed improves by 37%, and high-fidelity dynamic modeling is preserved even for 1080p-resolution videos.
📝 Abstract
In this paper, we investigate the challenge of spatio-temporal video prediction task, which involves generating future video frames based on historical spatio-temporal observation streams. Existing approaches typically utilize external information such as semantic maps to improve video prediction accuracy, which often neglect the inherent physical knowledge embedded within videos. Worse still, their high computational costs could impede their applications for high-resolution videos. To address these constraints, we introduce a novel framework called underline{P}hysics-underline{a}ssisted underline{S}patio-underline{t}emporal underline{Net}work (PastNet) for high-quality video prediction. The core of PastNet lies in incorporating a spectral convolution operator in the Fourier domain, which efficiently introduces inductive biases from the underlying physical laws. Additionally, we employ a memory bank with the estimated intrinsic dimensionality to discretize local features during the processing of complex spatio-temporal signals, thereby reducing computational costs and facilitating efficient high-resolution video prediction. Extensive experiments on various widely-used spatio-temporal video benchmarks demonstrate the effectiveness and efficiency of the proposed PastNet compared with a range of state-of-the-art methods, particularly in high-resolution scenarios.
Problem

Research questions and friction points this paper is trying to address.

Video Prediction
High Definition Videos
Physical Laws Integration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Physical Rule Integration
High Definition Video Prediction
Memory Optimization Technique
🔎 Similar Papers
2024-07-11Neural Information Processing SystemsCitations: 0
University of Science and Technology of China | Terminus Group | University of California, Los Angeles
H
Hao Wu
School of Computer Science and Technology, University of Science and Technology of China, Hefei, China
F
Fan Xu
School of Computer Science and Technology, University of Science and Technology of China, Hefei, China
C
Chong Chen
Terminus Group, Beijing, China
X
Xian-Sheng Hua
Terminus Group, Beijing, China
X
Xiao Luo
Department of Computer Science, University of California, Los Angeles, USA
Haixin Wang
Haixin Wang
UCLA; Peking University
AI for ScienceLarge Language ModelsMulti-modal LLMs