Score
Design and implement methods that extract patch- or grid-level features from spatiotemporal data and produce fine-grained localization of anomalies across both space and time. This involves learning region-level anomaly scores and jointly modeling spatial and temporal dependencies to generate spatiotemporal anomaly maps or patch-level anomaly heatmaps.
This work addresses the limitation of existing weakly supervised video anomaly detection methods, which primarily focus on temporal localization while lacking precise spatial awareness, thereby hindering interpretability in real-world applications. We propose a patch-based spatiotemporal anomaly localization framework that jointly models the temporal and spatial locations of anomalies using only video-level labels. Leveraging multiple instance learning, our approach infers region-level anomaly scores from grid-level patch features and introduces a novel neighborhood-aware Top-k spatiotemporal selection strategy to generate fine-grained spatial anomaly maps without requiring bounding box supervision. Additionally, we provide frame-level bounding box annotations for two widely used datasets. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art approaches across multiple benchmarks, achieving substantial improvements in spatiotemporal localization accuracy. The code, models, and new annotations are publicly released.
This work addresses the challenge in multivariate time series anomaly detection where over-generalized spatial structure modeling leads to erroneous reconstruction of anomalies and reduced recall. To mitigate this, we propose a prior-observation adversarial learning paradigm that jointly optimizes spatiotemporal dependency modeling. Our approach alternately learns an adjacency matrix as a structural prior in the spatial domain and employs a minimax adversarial mechanism to capture the discrepancy between this prior and data-driven observations, thereby significantly enhancing anomaly sensitivity along the temporal dimension and enabling, for the first time, channel-level anomaly localization. To facilitate systematic evaluation, we construct the first synthetic benchmark with precise channel-wise anomaly annotations. Extensive experiments demonstrate that our method achieves state-of-the-art performance on multiple public datasets as well as our newly introduced benchmark, excelling in both temporal detection and spatial localization tasks.
Existing multivariate time series anomaly detection methods often assume conditional independence among variables, neglecting time-varying, nonlinear spatiotemporal dependencies, and thus fail to identify “collective anomalies”—patterns where the system deviates anomalously as a whole despite individual variables appearing normal. To address this, we propose a decoupled modeling framework that separately learns temporal dynamics (via a Transformer encoder) and inter-variable spatial dependencies (via Copula-based multivariate likelihood modeling) in a shared latent space, jointly optimized through self-supervised contrastive learning. Our approach explicitly captures dynamic, nonlinear, and high-order dependencies without relying on strong independence assumptions. Evaluated on multiple real-world benchmarks, it achieves significant improvements over state-of-the-art methods—particularly in complex collective anomaly scenarios—demonstrating superior detection accuracy and robustness.
This work addresses the challenge of high annotation costs and the availability of only video-level weak labels in video anomaly detection by proposing a novel weakly supervised approach that, for the first time, jointly models spatiotemporal local anomalies under such constraints. The method treats normal and anomalous videos as negative and positive bags, respectively, within a multiple instance learning (MIL) framework. It integrates spatiotemporal feature extraction with a classifier-driven anomaly scoring mechanism and introduces a multiple instance ranking loss to effectively leverage video-level labels for pixel-level anomaly localization. Experimental results on the UCF Crime2Local dataset demonstrate that the proposed method accurately detects localized spatiotemporal anomalies using only video-level supervision.
Causal discovery in high-dimensional spatiotemporal grid data—common in climatology and neuroscience—is challenged by latent causal structures and strong spatial autocorrelation. Method: We propose the first framework integrating variational inference with structural causal modeling: (i) a reversible linear transformation ensures theoretical identifiability of latent time series and their causal relationships; (ii) built upon a variational autoencoder architecture, it jointly optimizes the evidence lower bound (ELBO) while incorporating spatiotemporal attention and sparse causal regularization for end-to-end causal graph learning. Contribution/Results: Our method significantly outperforms existing baselines on synthetic benchmarks, scales to grids with >10⁴ spatial units, and successfully recovers canonical teleconnection patterns—including the North Atlantic Oscillation (NAO) and Antarctic Oscillation (AAO)—from real-world climate data. It achieves both interpretability—via explicit causal graphs—and scalability—through efficient amortized inference—making it suitable for large-scale scientific discovery.
Weakly supervised video anomaly detection is often hindered by interference from background clutter and scene-level cues, leading to spatial bias and limited interpretability. To address this, this work proposes the SST-WSVAD framework, which dynamically sparsifies attention to focus on critical spatiotemporal regions and integrates end-to-end coupled spatial and temporal branches for fine-grained anomaly localization. The method innovatively introduces a motion-aware regularization mechanism that operates without external detectors or language prompts, enabling patch-level analysis of spatial bias and facilitating auditable evaluation. Evaluated on UCF-Crime, XD-Violence, and MSAD datasets, the approach achieves performance on par with state-of-the-art methods while offering fine-grained interpretability in both anomaly localization and scene-induced bias.
This work addresses the challenge of efficiently localizing an unknown number of anomalous regions in large-scale spatially dependent data. The authors propose SPLADE, a two-stage method that integrates intelligent sampling with boundary estimation to simultaneously and consistently estimate both the number and boundaries of multiple axis-aligned anomalous patches under general spatial dependence structures—without requiring full spatial grid segmentation. By leveraging a uniform Gaussian approximation and an efficient search strategy, SPLADE substantially improves computational efficiency and localization accuracy. Experimental results demonstrate that SPLADE outperforms existing approaches on both synthetic and real-world video surveillance datasets, achieving faster runtime, higher localization precision, and robustness to strong spatial dependencies.
Existing spatiotemporal forecasting methods suffer from a mismatch between model capacity and the inherent complexity of spatiotemporal dynamics, leading to performance bottlenecks and poor cross-domain generalization. This work proposes an adaptive dimensionality coordination framework that, for the first time, employs spatiotemporal entropy not as an optimization objective but as a diagnostic tool to identify such complexity mismatches. Guided by this insight, the framework dynamically balances spatial and temporal representations: it compresses spatial dimensions via low-rank matrix embeddings to preserve essential structural information while expanding the temporal horizon to capture long-range dependencies and mitigate error accumulation. Without increasing model capacity, the approach achieves substantial improvements in prediction accuracy and generalization across diverse domains—including traffic, meteorology, and epidemiology—demonstrating its broad applicability.
This work addresses the limitations of traditional spatiotemporal modeling approaches, which rely on covariance structures that tightly couple spatial and temporal components, leading to high computational costs. The authors propose a coarse-to-fine spatiotemporal modeling framework (CF-STM) that, for the first time, decouples multiscale locally weighted spatial representations from local state-space temporal models, enabling efficient, covariance-free modeling. This decoupling allows flexible selection of temporal dynamics without altering the spatial structure, substantially enhancing model interpretability and computational efficiency. Empirical evaluations demonstrate that CF-STM achieves predictive accuracy comparable to existing scalable methods in Monte Carlo simulations while incurring lower computational overhead. Furthermore, when applied to Tokyo residential land price data, the framework effectively captures complex spatiotemporal evolution patterns.