Score
Designs and implements models and signal‑processing pipelines that represent, encode, and extract features from data varying across space and time (spatio‑temporal data such as sequences of frames, spatial measurements over time, or streams of spatial tokens); this includes spatio‑temporal neural networks, filtering, spectral and clustering methods, and temporal factorization to produce compact representations. Analyzes temporal structure and dependencies — including temporal encoding schemes, boundary detection, and temporal feature learning — and builds efficient spatio‑temporal feature extractors and encoders for downstream tasks.
To address the task-specificity and poor generalizability of existing spatio-temporal deep learning models, this paper systematically surveys the full lifecycle of Spatio-Temporal Foundation Models (STFMs) and introduces the first structured pipeline framework. We propose a novel pipeline-oriented survey paradigm, establish a new taxonomy of data attributes tailored to spatio-temporal characteristics, and explicitly distinguish two core phases: “raw pretraining” and “downstream adaptation.” Within this framework, we unify the design principles for spatio-temporal embeddings, model architectures, pretraining objectives, and adaptation strategies. The framework supports data-driven modeling, multi-source dependency characterization, and multi-objective joint training—thereby significantly enhancing model reusability and development efficiency. It provides a reproducible methodological guide for STFM research and outlines concrete directions for future work.
Traditional deep learning models in spatiotemporal data science suffer from strong task specificity and heavy reliance on large-scale labeled datasets. Method: This paper formally introduces and defines the concept of Spatiotemporal Foundation Models (STFMs) for the first time, aiming to establish a general-purpose spatiotemporal intelligence paradigm applicable to urban computing, climate science, and related domains. It proposes a comprehensive STFM methodology taxonomy—integrating self-supervised pretraining, multimodal spatiotemporal representation learning, prompt engineering, and large-model architectural design—and identifies six core research directions. Contribution/Results: This work fills a critical gap by providing the first systematic survey and conceptual framework for STFMs. It advances both theoretical foundations and practical pathways toward spatiotemporal artificial general intelligence, offering essential guidance for future research and development in foundational spatiotemporal modeling.
This work addresses spatiotemporal prediction for high-resolution videos, aiming to efficiently and accurately synthesize future frames solely from historical frames and timestamps. We propose a physics-driven spectral-domain modeling framework: (1) physical priors are explicitly embedded into Fourier-frequency-domain convolutions to enforce physics-consistent inductive bias; and (2) a memory bank based on local intrinsic dimension estimation is introduced to enable feature discretization and compact storage. The resulting method is both lightweight and highly generalizable. Extensive evaluations on major video prediction benchmarks—including KTH, BAIR, and Moving MNIST—demonstrate substantial improvements over state-of-the-art methods: PSNR and SSIM increase by 1.2–2.8 dB, inference speed improves by 37%, and high-fidelity dynamic modeling is preserved even for 1080p-resolution videos.
Existing research on diffusion models for time-series and spatio-temporal data lacks a unified modeling framework and systematic task-specific analysis. Method: We propose the first multidimensional taxonomy tailored to temporal/spatio-temporal diffusion modeling, structured along three axes: model architecture (e.g., unconditional probabilistic vs. score-based models, conditional injection mechanisms), task paradigms (forecasting, generation, imputation, anomaly detection), and application domains (healthcare, climate, transportation, etc.). Contribution/Results: We construct a knowledge graph spanning six domains and deliver a reusable model selection guideline; explicitly characterize how conditional modeling enhances downstream task performance; and integrate key techniques—including diffusion process design, spatio-temporal graph neural networks, and multimodal fusion. This work establishes theoretical foundations and a technical roadmap for dynamic data modeling, advancing the paradigm shift of diffusion models toward time-series intelligence.
This work addresses critical challenges hindering the adoption of Spatio-Temporal Graph Neural Networks (ST-GNNs) in time-series classification and forecasting—namely, poor comparability, low reproducibility, limited interpretability, insufficient information capacity, and constrained scalability. To tackle these issues, we conduct a systematic literature review grounded in a structured meta-analysis of over 150 state-of-the-art studies. We propose the first cross-domain, unified benchmarking framework for horizontal comparison of ST-GNN models, systematically covering modeling paradigms, application scenarios, open-source implementations, benchmark datasets, and evaluation metrics. Furthermore, we integrate models, code, data, and empirical results into the first open, reusable ST-GNN knowledge graph. Finally, we provide standardized evaluation guidelines and concrete improvement pathways. This synthesis establishes a rigorous, transparent foundation for both methodological innovation and empirical validation in ST-GNN research.
Real-world spatiotemporal data arrive in streams, and the underlying graph structure dynamically expands—posing dual challenges for online forecasting: inefficient model retraining and catastrophic forgetting. To address this, we propose the first prompt-tuning framework for dynamic spatiotemporal graphs, grounded in two principles—*expansion* and *compression*. Our method employs a reusable, continuous prompt pool to retain historical knowledge while enabling rapid adaptation to newly added nodes or time intervals. It integrates lightweight prompt fine-tuning of spatiotemporal graph neural networks, dynamic prompt pool management, and joint optimization. Evaluated on multiple real-world traffic and environmental datasets, our approach consistently outperforms state-of-the-art methods. Crucially, it achieves superior prediction accuracy, training efficiency, and cross-scenario generalization while introducing fewer than 0.5% additional parameters.
Existing model evaluation methods for spatiotemporal data—characterized by co-occurring missingness and heterogeneity, strong nonlinearity, and nonstationarity—lack interpretability and robustness. Method: We propose the first assumption-free, distribution-agnostic residual correlation diagnostic framework. It quantifies residual dependence structures across spatiotemporal dimensions via spatiotemporal graph modeling and asymptotically distribution-free autocorrelation statistics, enabling precise localization of local underfitting regions. Crucially, it imposes no prior assumptions on data distribution or underlying dynamics and natively supports interpretability assessment for sparse observations and nonlinear models—including spatiotemporal graph neural networks. Results: Extensive validation on synthetic and real-world datasets demonstrates that our framework accurately identifies performance-weak subregions, significantly enhancing the targeting and efficiency of model iteration.
This work addresses the longstanding trade-off in spatiotemporal modeling: convolutional neural networks (CNNs) struggle to capture long-range dependencies, while Transformer-based approaches incur prohibitive computational costs. To reconcile this dilemma, the authors propose MIMO-ESP, a purely CNN-based architecture that achieves Transformer-like global receptive fields through a multi-input multi-output design, decoupled temporal modeling, and dilated convolutions. This approach effectively integrates spatial and temporal information while preserving the parallelizability and efficiency inherent to CNNs, substantially reducing computational complexity. Extensive experiments demonstrate that MIMO-ESP consistently outperforms state-of-the-art methods across three diverse benchmark datasets—covering video prediction, traffic flow forecasting, and precipitation nowcasting—thereby achieving an optimal balance between predictive accuracy and computational efficiency.
This work proposes STemDist, the first dual-dimensional dataset distillation framework tailored for spatiotemporal forecasting. Existing methods typically compress only a single dimension—either temporal or spatial—limiting their efficiency in large-scale spatiotemporal sequence prediction. In contrast, STemDist jointly compresses both time and space dimensions by integrating cluster-level coarse-grained distillation with subset-level fine-grained optimization. This approach significantly reduces training overhead while simultaneously improving prediction accuracy. Extensive experiments on five real-world datasets demonstrate that STemDist achieves up to 6× faster training, 8× less memory consumption, and up to a 12% reduction in prediction error compared to state-of-the-art baselines.
This work proposes ST-Prune, a novel approach that introduces a learning-complexity-based dynamic sample pruning mechanism into spatiotemporal forecasting. Addressing the inefficiency of conventional training paradigms that rely on redundant static datasets, ST-Prune adaptively selects high-informativeness samples by continuously evaluating the model’s learning state during training. This method breaks away from the traditional static data iteration framework, enabling significantly accelerated training across multiple real-world spatiotemporal datasets while maintaining or even improving predictive performance. The approach demonstrates strong generalizability and scalability, offering a more efficient and effective alternative for training spatiotemporal prediction models.
Existing studies struggle to disentangle the functional role of membrane potential propagation in complex temporal modeling within spiking neural networks (SNNs), while tight coupling between spatial semantics and temporal dynamics induces spatiotemporal resource competition. To address this, we propose STSep—a novel decoupled architecture that, for the first time in SNNs, separates spatial and temporal processing branches. Through state-free membrane potential analysis, we discover that moderate suppression of membrane potential propagation enhances performance, revealing an intrinsic trade-off in spatiotemporal encoding. We further design a spatial-temporal separable residual block that integrates explicit temporal differencing with spike-based attention for efficient dynamic modeling. STSep achieves state-of-the-art accuracy on Something-Something V2, UCF101, and HMDB51. Retrieval experiments and attention visualization confirm its strong focus on motion-related features, significantly outperforming static appearance-based approaches.