Score
Designs and implements models, feature‑extraction pipelines, and analytic methods that operate on sequences of per‑token probability estimates (probability signals), modeling their temporal and spectral structure to predict, detect, or characterize events. This includes time‑series and time‑frequency feature extraction, sequence modeling architectures, and robustness techniques for detection or classification across distribution shifts.
Long-sequence modeling faces fundamental challenges including limited context length, difficulty in capturing long-range dependencies, and low efficiency in online learning. To address these, this work systematically reviews the resurgence of state-space models (SSMs) and recurrent computation, proposing a novel local forward-computation paradigm tailored for real-world online learning—thereby circumventing the temporal backtracking constraints inherent in standard backpropagation through time (BPTT). We introduce the first unified taxonomy encompassing both deep SSMs and large-context Transformers. Our framework integrates structured linear attention, enhanced RNN architectures, local recurrence mechanisms, and online optimization algorithms. The study rigorously clarifies the theoretical representational advantages and practical sequential reasoning benefits of recurrent modeling over alternatives. Collectively, this work delivers a scalable technical roadmap for low-latency, highly extensible long-sequence modeling.
Classical nonlinear filtering theory lacks integration with modern Transformer architectures, and existing sequence modeling approaches suffer from insufficient interpretability of underlying dynamical mechanisms in probabilistic modeling and sequential inference. Method: We model each Transformer layer’s activations as conditional distribution surrogates of the posterior measure under a hidden Markov model (HMM), interpret self-attention as a nonlinear Bayesian predictor, and formalize inter-layer transformations as fixed-point iterations on the space of probability measures. Contribution/Results: We establish, for the first time, a rigorous correspondence between Transformer forward propagation and nonlinear filtering—specifically particle filtering and Gaussian filtering—deriving explicit fixed-point update formulas in the HMM special case. This framework endows attention mechanisms with a principled probabilistic semantics and provides a theoretical foundation for designing interpretable, robust inference architectures grounded in stochastic filtering theory.
The time-series foundation model (TSFM) field lacks a systematic taxonomy, hindering comparative analysis and principled design. Method: We propose the first multi-dimensional taxonomy tailored to Transformer-based TSFMs, spanning five dimensions: architectural design, forecasting paradigm, variable dimensionality, scale/complexity, and pretraining objective functions—introducing objective function type as a novel classification criterion to unify capability characterization and design rationale. Through comprehensive literature review, architectural analysis, task mapping, and paradigm comparison, we systematically cover mainstream modeling approaches—including patch-based and raw-sequence methods. Contribution/Results: This taxonomy establishes a structured knowledge graph for TSFMs, clarifying technological trends, exposing critical research gaps, and providing a principled foundation for developing scalable, interpretable, and multi-task-cooperative time-series foundation models.
Time-series forecasting suffers from poor generalization and low efficiency due to tight coupling among sequence representation, information extraction, and future projection. To address this, we propose a modular forecasting framework that decouples the pipeline into three independently optimizable stages—representation learning, information extraction, and target projection—enabling flexible, task-aware component configuration. Our approach innovatively integrates convolutional layers with a lightweight self-attention mechanism, achieving efficient local feature modeling while capturing long-range temporal dependencies. Evaluated on seven benchmark datasets, the method consistently outperforms existing state-of-the-art models in prediction accuracy, while requiring significantly fewer parameters and achieving faster training and inference speeds. This demonstrates substantial improvements in both statistical performance and computational efficiency.
This study addresses the challenge of inefficient inference in Gaussian processes for sequential signal processing by moving beyond the conventional machine learning reliance on independent and identically distributed assumptions. It proposes a unified sequential inference framework for Gaussian processes, systematically integrating techniques from sequential Bayesian inference, incremental learning, streaming computation, and state-space modeling, with signal processing as the central organizing principle. This work not only bridges the longstanding gap between modern machine learning and classical signal processing but also delivers scalable and efficient practical solutions—along with a clear deployment roadmap—for time-series forecasting, anomaly detection, adaptive sensing, and real-time Bayesian optimization.
This paper addresses the core challenge in change detection for nonstationary processes—where the post-change distribution evolves over time and is a priori unknown. Method: We propose a robust sequential detection framework based on the least favorable distribution, integrating minimax decision theory, nonstationary statistical modeling, and sequential hypothesis testing to construct an adaptive CUSUM-type algorithm capable of real-time, robust response to dynamic distributional shifts. Contribution/Results: To our knowledge, this is the first work to establish asymptotically minimax-optimal theoretical guarantees for sequential change detection under nonstationarity. Empirical evaluation on real-world public health and military monitoring datasets, as well as extensive simulations, demonstrates that our method reduces average detection delay by 23% compared to conventional approaches, while maintaining stable and controllable false alarm rates—thereby validating both theoretical soundness and practical superiority.
This study addresses the challenges of modeling and analyzing complex time series arising in astrophysics, meteorology, finance, and other domains by systematically integrating classical statistical methods—such as ARIMA, exponential smoothing, and state-space models—with modern machine learning techniques, including tree-based ensembles, hidden Markov models, Gaussian processes, and deep learning architectures like RNNs, CNNs, and Transformers. By distilling cross-disciplinary modeling principles, the work establishes a unified framework that combines theoretical rigor with practical guidance, offering researchers a comprehensive and extensible toolkit for time series analysis. This approach significantly enhances the capacity to handle temporal data across diverse scientific and applied contexts.
Existing approaches to modeling the joint probability distribution of irregular multivariate time series often struggle to simultaneously maintain model expressiveness and marginal consistency, leading to unreliable or contradictory predictions. This work proposes CircuITS, the first method to introduce probabilistic circuits into this task, constructing a structured, interpretable joint probabilistic model that supports efficient inference. By leveraging the compositional structure of probabilistic circuits, CircuITS flexibly captures complex dependencies among variables while rigorously preserving consistency between the joint distribution and its marginals. Empirical evaluation on four real-world datasets demonstrates that CircuITS significantly outperforms state-of-the-art methods in both joint and marginal density estimation.
This work addresses the limitations of conventional Transformers in time series modeling—namely, their lack of interpretability, inadequate channel-wise modeling capacity, and absence of structured conditional generation mechanisms. The authors propose the first systematic extension of Probabilistic Transformers (PT) to Spatio-Temporal Probabilistic Transformers (ST-PT), formulated as programmable factor graphs. By explicitly designing graph topology, conditional potential functions, and a message-passing mechanism driven by variational inference, ST-PT achieves three key innovations: incorporating structural priors to enhance few-shot performance, enabling external conditions to structurally govern the generative process, and integrating CRF-based teacher distillation of latent variables to mitigate autoregressive error accumulation. Experimental results demonstrate that ST-PT serves as a flexible and effective general-purpose backbone for spatio-temporal modeling.
For large-scale discrete-event systems, analyzing state-transition event time series (STE-ts) suffers from reliance on distortion-prone resampling and difficulty in jointly modeling transition timing and state dwell duration. To address this, we propose the Selective Time Hamming (STH) distance, which directly encodes event occurrence times and state residence durations without resampling. STH unifies Hamming distance and Jaccard similarity within a single metric and supports multi-state focused matching. It preserves computational efficiency while significantly improving pattern recognition accuracy. Experiments on synthetic and real-world datasets demonstrate that STH achieves an average 2.3× speedup over baseline methods and improves clustering and anomaly detection F1-scores by 12.6%.
This study systematically investigates the frequency-domain encoding capabilities of the Chronos foundation model, addressing a critical gap in understanding how such models represent fundamental signal properties. Through controlled experiments using discrete sinusoidal signals and a lightweight online Minimum Description Length (MDL) probing framework, the work examines the existence, separability, and cross-spectral fidelity of internal frequency representations within the Chronos decoder. The research reveals, for the first time, a degradation in representation quality in high-frequency regions, thereby delineating both the strengths and limitations of Chronos’s frequency encoding mechanism. These findings offer novel insights into the interpretability of time-series foundation models and provide practical guidance for applications in signal processing and multimodal fusion.