Score
Design and implement encoders that represent localized spatial structure across predefined regions and the temporal dynamics of signals over time, using region-based graph constructions and spatiotemporal modeling (e.g., graph convolutions combined with temporal convolutions or attention). Build representations and modules that capture region-wise interactions and temporal dependencies while remaining robust to variations in sensor/channel configuration.
To address the task-specificity and poor generalizability of existing spatio-temporal deep learning models, this paper systematically surveys the full lifecycle of Spatio-Temporal Foundation Models (STFMs) and introduces the first structured pipeline framework. We propose a novel pipeline-oriented survey paradigm, establish a new taxonomy of data attributes tailored to spatio-temporal characteristics, and explicitly distinguish two core phases: “raw pretraining” and “downstream adaptation.” Within this framework, we unify the design principles for spatio-temporal embeddings, model architectures, pretraining objectives, and adaptation strategies. The framework supports data-driven modeling, multi-source dependency characterization, and multi-objective joint training—thereby significantly enhancing model reusability and development efficiency. It provides a reproducible methodological guide for STFM research and outlines concrete directions for future work.
Traditional deep learning models in spatiotemporal data science suffer from strong task specificity and heavy reliance on large-scale labeled datasets. Method: This paper formally introduces and defines the concept of Spatiotemporal Foundation Models (STFMs) for the first time, aiming to establish a general-purpose spatiotemporal intelligence paradigm applicable to urban computing, climate science, and related domains. It proposes a comprehensive STFM methodology taxonomy—integrating self-supervised pretraining, multimodal spatiotemporal representation learning, prompt engineering, and large-model architectural design—and identifies six core research directions. Contribution/Results: This work fills a critical gap by providing the first systematic survey and conceptual framework for STFMs. It advances both theoretical foundations and practical pathways toward spatiotemporal artificial general intelligence, offering essential guidance for future research and development in foundational spatiotemporal modeling.
This work addresses critical challenges hindering the adoption of Spatio-Temporal Graph Neural Networks (ST-GNNs) in time-series classification and forecasting—namely, poor comparability, low reproducibility, limited interpretability, insufficient information capacity, and constrained scalability. To tackle these issues, we conduct a systematic literature review grounded in a structured meta-analysis of over 150 state-of-the-art studies. We propose the first cross-domain, unified benchmarking framework for horizontal comparison of ST-GNN models, systematically covering modeling paradigms, application scenarios, open-source implementations, benchmark datasets, and evaluation metrics. Furthermore, we integrate models, code, data, and empirical results into the first open, reusable ST-GNN knowledge graph. Finally, we provide standardized evaluation guidelines and concrete improvement pathways. This synthesis establishes a rigorous, transparent foundation for both methodological innovation and empirical validation in ST-GNN research.
Traditional ST-GCNs employ单一 temporal modules—either CNNs or LSTMs—leading to insufficient capture of dynamic spatiotemporal patterns. To address this, we propose a plug-and-play hybrid temporal module that, for the first time, synergistically integrates CNNs and LSTMs within a unified co-temporal block. This design jointly models local temporal features and long-range dependencies. Through theoretical analysis and cross-dataset ablation studies, we systematically characterize the intrinsic relationship between temporal module architecture and representational capacity. Evaluated on standard spatiotemporal graph benchmarks—including NTU-RGB+D and PeMSD7—our method achieves significant improvements in prediction accuracy and cross-domain generalization. It consistently outperforms pure-CNN and pure-LSTM baselines in temporal representation learning. The proposed module establishes a reusable, principled design paradigm for temporal modeling in ST-GCNs, advancing both expressiveness and architectural flexibility.
To address poor robustness in regional representation caused by strong noise and sparse labels in urban spatiotemporal graph data, this paper introduces the Spatiotemporal Heterogeneous Graph Neural Encoder (ST-HGAE), the first application of masked autoencoding to spatiotemporal graph learning. ST-HGAE jointly masks node features and graph structure, enabling generative self-supervised learning to automatically distill dynamic spatiotemporal dependencies. Its core innovations include a structure-aware masking strategy tailored for heterogeneous spatiotemporal graphs, and a dual reconstruction objective integrating node-level feature recovery with topology reconstruction. Evaluated on traffic flow, pedestrian flow, and crime prediction tasks, ST-HGAE consistently outperforms state-of-the-art methods—particularly under high noise levels and low label rates—while significantly enhancing modeling of dynamic spatial correlations among regions.
Existing grid-based graph neural networks (GNNs) are constrained by first-order node-edge message passing, limiting their ability to capture high-order spatial dependencies among regional features and voxel-level geometric information in physical systems. Method: We propose the Cell-embedded Graph Neural Network (CeGNN), which introduces a novel learnable cell-attribute embedding mechanism that extends message passing from “edge→node” to “cell→edge→node”, explicitly encoding local geometric structure; additionally, we design a basis-function-driven feature enhancement module that mitigates over-smoothing by leveraging implicit features as functional bases. Contribution/Results: Evaluated on diverse PDE-solving tasks and real-world physical datasets, CeGNN consistently outperforms state-of-the-art GNN baselines, achieving up to an order-of-magnitude reduction in prediction error. This work establishes a new paradigm for data-driven physics simulation—uniquely balancing expressive power with geometric awareness.
Existing spatiotemporal graph models—such as GNNs and Transformers—typically adopt decoupled spatial and temporal modeling, limiting their ability to capture nontrivial structural dependencies inherent in graph topologies. To address this, we propose Cy2Mixer, a three-module architecture integrating temporal modeling, standard message passing, and a novel recurrent message-passing block (RMPB). The RMPB’s core innovation lies in explicitly constructing and aggregating cycle subgraphs to encode topological invariants; we theoretically prove its information complementarity with conventional message passing, thereby overcoming the limitations of spatiotemporal decoupling. Additionally, Cy2Mixer incorporates a gMLP backbone, gating mechanisms, and mathematically grounded topological representations. Extensive experiments on multiple spatiotemporal forecasting benchmarks demonstrate state-of-the-art performance, with significant improvements in prediction accuracy and generalization—particularly for traffic flow forecasting.
Existing model evaluation methods for spatiotemporal data—characterized by co-occurring missingness and heterogeneity, strong nonlinearity, and nonstationarity—lack interpretability and robustness. Method: We propose the first assumption-free, distribution-agnostic residual correlation diagnostic framework. It quantifies residual dependence structures across spatiotemporal dimensions via spatiotemporal graph modeling and asymptotically distribution-free autocorrelation statistics, enabling precise localization of local underfitting regions. Crucially, it imposes no prior assumptions on data distribution or underlying dynamics and natively supports interpretability assessment for sparse observations and nonlinear models—including spatiotemporal graph neural networks. Results: Extensive validation on synthetic and real-world datasets demonstrates that our framework accurately identifies performance-weak subregions, significantly enhancing the targeting and efficiency of model iteration.
Existing graph neural networks are inherently limited to modeling pairwise relationships, struggling to effectively capture higher-order topological structures while suffering from rapidly escalating computational complexity as graph size grows. To address these limitations, this work proposes a simplicial complex–based spatiotemporal neural network that, for the first time, integrates simplicial complexes into spatiotemporal modeling. By leveraging spatiotemporal random walks on high-dimensional simplicial complexes and parallelized temporal convolutions, the proposed method transcends the pairwise interaction constraints of conventional graph neural networks. This approach significantly enhances the capacity to model higher-order topological dependencies in complex systems while maintaining computational efficiency.
Existing spatial semantic representations struggle to effectively reason about structured temporal dynamics—such as the periodic movement of household objects—in semi-static environments. This work proposes PredictiveGraphs, a predictive 3D scene graph that integrates spatiotemporal and semantic information by embedding Perpetua* Bayesian filters directly into inter-node relationships, enabling temporal modeling and future prediction of object states. By jointly modeling spatiotemporal-semantic relations and performing recurrent state inference, the approach maintains robustness under distributional shifts. Evaluated over three-week navigation tasks in both simulation and real-world settings—with environmental changes occurring every two hours—the method significantly outperforms current baselines in accurately forecasting the dynamic evolution of the environment.
This study investigates the internal mechanisms underlying spatiotemporal reasoning in vision-language models (VLMs), with a focus on how spatial and textual representations are integrated. Through causal interventions, linear probing, representational analysis, and cross-modal activation alignment, the work systematically demonstrates the widespread presence of linear spatial and temporal identifiers in both image and video VLMs. It further reveals, for the first time, that these mechanisms modulate belief states in intermediate model layers. This insight not only offers a novel perspective for improving VLM interpretability and alignment design but also serves as a diagnostic tool to uncover model limitations and generate informative training signals.
Existing 4D generation methods struggle to simultaneously ensure global appearance consistency and local dynamic realism, often lacking deep modeling of physical spatiotemporal regularities. This work proposes the first generative framework embedded with a 4D spatiotemporal cognitive mechanism: it constructs global appearance and local dynamic graphs from multimodal features, employs a semantic bridging fusion strategy to form a unified 4D cognitive graph, and integrates a world model to reason future states, thereby guiding a latent diffusion model to generate structurally plausible and topologically consistent 4D Gaussian scenes. Evaluated on both a newly curated and integrated ST-4D dataset, the method significantly improves geometric plausibility and spatiotemporal coherence in both 3D and 4D generation tasks.