token-probability sequence modeling

Designs and implements models, feature‑extraction pipelines, and analytic methods that operate on sequences of per‑token probability estimates (probability signals), modeling their temporal and spectral structure to predict, detect, or characterize events. This includes time‑series and time‑frequency feature extraction, sequence modeling architectures, and robustness techniques for detection or classification across distribution shifts.

token-probabilitysequencemodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.05
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Classical nonlinear filtering theory lacks integration with modern Transformer architectures, and existing sequence modeling approaches suffer from insufficient interpretability of underlying dynamical mechanisms in probabilistic modeling and sequential inference. Method: We model each Transformer layer’s activations as conditional distribution surrogates of the posterior measure under a hidden Markov model (HMM), interpret self-attention as a nonlinear Bayesian predictor, and formalize inter-layer transformations as fixed-point iterations on the space of probability measures. Contribution/Results: We establish, for the first time, a rigorous correspondence between Transformer forward propagation and nonlinear filtering—specifically particle filtering and Gaussian filtering—deriving explicit fixed-point update formulas in the HMM special case. This framework endows attention mechanisms with a principled probabilistic semantics and provides a theoretical foundation for designing interpretable, robust inference architectures grounded in stochastic filtering theory.

Bridging classical nonlinear filtering with modern inferenceDescribing layer operations as fixed-point updatesInterpreting transformer signals as conditional measure surrogates

Foundation Models for Time Series: A Survey

Apr 05, 2025
SR
Siva Rama Krishna Kottapalli
🏛️ Dell Technologies | University of Massachusetts Lowell | Worcester Polytechnic Institute

The time-series foundation model (TSFM) field lacks a systematic taxonomy, hindering comparative analysis and principled design. Method: We propose the first multi-dimensional taxonomy tailored to Transformer-based TSFMs, spanning five dimensions: architectural design, forecasting paradigm, variable dimensionality, scale/complexity, and pretraining objective functions—introducing objective function type as a novel classification criterion to unify capability characterization and design rationale. Through comprehensive literature review, architectural analysis, task mapping, and paradigm comparison, we systematically cover mainstream modeling approaches—including patch-based and raw-sequence methods. Contribution/Results: This taxonomy establishes a structured knowledge graph for TSFMs, clarifying technological trends, exposing critical research gaps, and providing a principled foundation for developing scalable, interpretable, and multi-task-cooperative time-series foundation models.

Classifying models by architecture, prediction type, and data handlingIdentifying trends and future research in time series modelingSurveying transformer-based foundation models for time series analysis

Decomposing the Time Series Forecasting Pipeline: A Modular Approach for Time Series Representation, Information Extraction, and Projection

Jul 08, 2025
RL
Robert Leppich
🏛️ University of Wuerzburg | Illinois Institute of Technology

Time-series forecasting suffers from poor generalization and low efficiency due to tight coupling among sequence representation, information extraction, and future projection. To address this, we propose a modular forecasting framework that decouples the pipeline into three independently optimizable stages—representation learning, information extraction, and target projection—enabling flexible, task-aware component configuration. Our approach innovatively integrates convolutional layers with a lightweight self-attention mechanism, achieving efficient local feature modeling while capturing long-range temporal dependencies. Evaluated on seven benchmark datasets, the method consistently outperforms existing state-of-the-art models in prediction accuracy, while requiring significantly fewer parameters and achieving faster training and inference speeds. This demonstrates substantial improvements in both statistical performance and computational efficiency.

Effective sequence representation for time series forecastingMeaningful information extraction from time series dataPrecise future projection in diverse forecasting tasks

This study addresses the challenge of inefficient inference in Gaussian processes for sequential signal processing by moving beyond the conventional machine learning reliance on independent and identically distributed assumptions. It proposes a unified sequential inference framework for Gaussian processes, systematically integrating techniques from sequential Bayesian inference, incremental learning, streaming computation, and state-space modeling, with signal processing as the central organizing principle. This work not only bridges the longstanding gap between modern machine learning and classical signal processing but also delivers scalable and efficient practical solutions—along with a clear deployment roadmap—for time-series forecasting, anomaly detection, adaptive sensing, and real-time Bayesian optimization.

Gaussian ProcessesSequential InferenceSignal Processing

Robust quickest change detection in nonstationary processes

Oct 14, 2023
YH
Yingze Hou
🏛️ University of Pittsburgh | Indian Space Research Organization | George Mason University

This paper addresses the core challenge in change detection for nonstationary processes—where the post-change distribution evolves over time and is a priori unknown. Method: We propose a robust sequential detection framework based on the least favorable distribution, integrating minimax decision theory, nonstationary statistical modeling, and sequential hypothesis testing to construct an adaptive CUSUM-type algorithm capable of real-time, robust response to dynamic distributional shifts. Contribution/Results: To our knowledge, this is the first work to establish asymptotically minimax-optimal theoretical guarantees for sequential change detection under nonstationarity. Empirical evaluation on real-world public health and military monitoring datasets, as well as extensive simulations, demonstrates that our method reduces average detection delay by 23% compared to conventional approaches, while maintaining stable and controllable false alarm rates—thereby validating both theoretical soundness and practical superiority.

Apply algorithms to health and military monitoringDetect changes in non-stationary processes robustlyHandle unknown post-change distribution variations

Latest Papers

What's happening recently
View more

This study addresses the challenges of modeling and analyzing complex time series arising in astrophysics, meteorology, finance, and other domains by systematically integrating classical statistical methods—such as ARIMA, exponential smoothing, and state-space models—with modern machine learning techniques, including tree-based ensembles, hidden Markov models, Gaussian processes, and deep learning architectures like RNNs, CNNs, and Transformers. By distilling cross-disciplinary modeling principles, the work establishes a unified framework that combines theoretical rigor with practical guidance, offering researchers a comprehensive and extensible toolkit for time series analysis. This approach significantly enhances the capacity to handle temporal data across diverse scientific and applied contexts.

ForecastingMachine LearningStatistical Modeling

Existing approaches to modeling the joint probability distribution of irregular multivariate time series often struggle to simultaneously maintain model expressiveness and marginal consistency, leading to unreliable or contradictory predictions. This work proposes CircuITS, the first method to introduce probabilistic circuits into this task, constructing a structured, interpretable joint probabilistic model that supports efficient inference. By leveraging the compositional structure of probabilistic circuits, CircuITS flexibly captures complex dependencies among variables while rigorously preserving consistency between the joint distribution and its marginals. Empirical evaluation on four real-world datasets demonstrates that CircuITS significantly outperforms state-of-the-art methods in both joint and marginal density estimation.

consistent marginalizationirregular multivariate time seriesjoint probabilistic modeling

This work addresses the limitations of conventional Transformers in time series modeling—namely, their lack of interpretability, inadequate channel-wise modeling capacity, and absence of structured conditional generation mechanisms. The authors propose the first systematic extension of Probabilistic Transformers (PT) to Spatio-Temporal Probabilistic Transformers (ST-PT), formulated as programmable factor graphs. By explicitly designing graph topology, conditional potential functions, and a message-passing mechanism driven by variational inference, ST-PT achieves three key innovations: incorporating structural priors to enhance few-shot performance, enabling external conditions to structurally govern the generative process, and integrating CRF-based teacher distillation of latent variables to mitigate autoregressive error accumulation. Experimental results demonstrate that ST-PT serves as a flexible and effective general-purpose backbone for spatio-temporal modeling.

Conditional Random FieldFactor GraphProbabilistic Transformer

A Selective Temporal Hamming distance to find patterns in state transition event timeseries, at scale

Dec 01, 2025
SM
Sylvain Marié
🏛️ AI Hub | Schneider Electric | IntenCity

For large-scale discrete-event systems, analyzing state-transition event time series (STE-ts) suffers from reliance on distortion-prone resampling and difficulty in jointly modeling transition timing and state dwell duration. To address this, we propose the Selective Time Hamming (STH) distance, which directly encodes event occurrence times and state residence durations without resampling. STH unifies Hamming distance and Jaccard similarity within a single metric and supports multi-state focused matching. It preserves computational efficiency while significantly improving pattern recognition accuracy. Experiments on synthetic and real-world datasets demonstrate that STH achieves an average 2.3× speedup over baseline methods and improves clustering and anomaly detection F1-scores by 12.6%.

Avoids distorting resampling in large event sequence databasesDevelops a distance metric for state transition event timeseriesGeneralizes Hamming and Jaccard metrics with improved precision

This study systematically investigates the frequency-domain encoding capabilities of the Chronos foundation model, addressing a critical gap in understanding how such models represent fundamental signal properties. Through controlled experiments using discrete sinusoidal signals and a lightweight online Minimum Description Length (MDL) probing framework, the work examines the existence, separability, and cross-spectral fidelity of internal frequency representations within the Chronos decoder. The research reveals, for the first time, a degradation in representation quality in high-frequency regions, thereby delineating both the strengths and limitations of Chronos’s frequency encoding mechanism. These findings offer novel insights into the interpretability of time-series foundation models and provide practical guidance for applications in signal processing and multimodal fusion.

foundation modelsfrequency representationmodel interpretability

Hot Scholars

SM

Siwei Ma

Peking University
Video Coding and Processing
KM

Krikamol Muandet

CISPA - Helmholtz Center for Information Security
Kernel MethodsCounterfactual LearningStatistical Learning TheoryMachine Learning
MC

Michele Caprio

Lecturer (Asst. Prof.), The University of Manchester
Imprecise ProbabilityApplied ProbabilityArtificial IntelligenceStatistical Theory
MI

Michael I. Jordan

Professor of Electrical Engineering and Computer Sciences and Professor of Statistics, UC Berkeley
machine learningcomputer sciencestatisticsartificial intelligence
KD

Kevin Duh

Johns Hopkins University
Natural Language ProcessingMachine Learning