Score
Designs and implements models and architectures that represent, analyze, and predict time-varying patterns and dependencies in sequential or event-based data, covering sequence models, latent-state dynamics, duration and time-to-event processes, temporal decay and difference effects, and temporal coherence/context. This work builds and evaluates recurrent, convolutional, transformer and other temporal-modeling architectures and techniques to capture short- and long-range temporal correlations, motion, timing, and temporal prediction for forecasting, classification, and event-risk estimation.
Long-sequence modeling faces fundamental challenges including limited context length, difficulty in capturing long-range dependencies, and low efficiency in online learning. To address these, this work systematically reviews the resurgence of state-space models (SSMs) and recurrent computation, proposing a novel local forward-computation paradigm tailored for real-world online learning—thereby circumventing the temporal backtracking constraints inherent in standard backpropagation through time (BPTT). We introduce the first unified taxonomy encompassing both deep SSMs and large-context Transformers. Our framework integrates structured linear attention, enhanced RNN architectures, local recurrence mechanisms, and online optimization algorithms. The study rigorously clarifies the theoretical representational advantages and practical sequential reasoning benefits of recurrent modeling over alternatives. Collectively, this work delivers a scalable technical roadmap for low-latency, highly extensible long-sequence modeling.
Large language models (LLMs) inherently struggle to capture continuous temporal dynamics and explicit inter-variable dependencies in time series analysis. Method: This paper systematically reviews the emerging “time-series-to-image + vision model” paradigm, proposing the first dual-dimensional taxonomy: (i) time-series image encoding strategies (e.g., Gramian Angular Field, Markov Transition Field) and (ii) vision-model adaptation architectures (e.g., Vision Transformers, multimodal alignment, feature-decoupled reconstruction). It rigorously defines key pre-/post-processing challenges, surveys over 100 works, and establishes a unified evaluation framework. Contribution/Results: Empirical results demonstrate that vision-based approaches consistently outperform pure sequence models—achieving average accuracy gains of 5–12% across anomaly detection, forecasting, and classification tasks—thereby offering a promising new direction for time-series modeling.
The time-series foundation model (TSFM) field lacks a systematic taxonomy, hindering comparative analysis and principled design. Method: We propose the first multi-dimensional taxonomy tailored to Transformer-based TSFMs, spanning five dimensions: architectural design, forecasting paradigm, variable dimensionality, scale/complexity, and pretraining objective functions—introducing objective function type as a novel classification criterion to unify capability characterization and design rationale. Through comprehensive literature review, architectural analysis, task mapping, and paradigm comparison, we systematically cover mainstream modeling approaches—including patch-based and raw-sequence methods. Contribution/Results: This taxonomy establishes a structured knowledge graph for TSFMs, clarifying technological trends, exposing critical research gaps, and providing a principled foundation for developing scalable, interpretable, and multi-task-cooperative time-series foundation models.
Existing sequence models often neglect explicit temporal information, limiting their capacity for temporal reasoning and event timing modeling. This work proposes ChronoSSM, an autoregressive state space model that jointly models event content and timestamps within the SSM framework for the first time. By employing a shared backbone network, ChronoSSM simultaneously optimizes event prediction and time generation objectives, departing from conventional two-stage paradigms and enabling learned representations to encapsulate both semantic and temporal structures. Experiments across four datasets with varying levels of temporal annotation density demonstrate that ChronoSSM substantially improves the recoverability of inter-event timing information while maintaining strong event generation quality.
Real-world multivariate time series exhibit strong non-stationarity and cross-scale dynamics; however, prevailing models rely on fixed-scale priors—such as chunk-based tokenization or static frequency-domain transformations—limiting modeling flexibility and hindering robustness to abrupt, high-magnitude events. To address this, we propose an adaptive hierarchical architecture that jointly captures instantaneous fluctuations and long-term trends via multi-scale convolutional encoding, integrates sequential modeling using BiLSTM or Transformer backbones, incorporates Squeeze-and-Excitation gating for channel-wise feature recalibration, and employs multi-head temporal attention for context-aware dynamic feature fusion. The resulting framework establishes a unified paradigm for time-series modeling. Evaluated across 32 benchmark datasets on forecasting, imputation, and classification tasks, our method achieves state-of-the-art performance on 24 datasets—outperforming leading approaches including EMTSF, TimesNet, and PatchTST by significant margins.
Existing deep learning architectures for time series forecasting—ranging from MLPs, CNNs, RNNs, GNNs, Transformers, and Mamba to diffusion models and foundation models—face persistent challenges in modeling channel dependencies, ensuring distributional shift robustness, enabling causal inference, and achieving feature disentanglement. Method: We propose an “architectural renaissance” perspective, empirically revealing the surprising competitiveness of linear models across diverse scenarios; establish a unified evaluation framework grounded in intrinsic data characteristics, systematically mapping architecture choices to task properties for the first time; and integrate hybrid modeling, generative modeling, and mechanism-driven approaches into a cross-paradigm unification pathway. Contribution/Results: Our work delivers an open problem inventory, standardized evaluation dimensions, and low-barrier practical guidelines—shifting time series intelligence from heuristic “architecture stacking” toward principled “mechanism-aware adaptation.”
This paper addresses the problem of predicting the distribution of future events in human action sequences—emphasizing the composition of likely outcomes rather than precise temporal ordering—a task critical for applications in retail, finance, healthcare, and recommender systems. To overcome the limitations of dominant autoregressive paradigms—namely, their strong sequential dependency and tendency toward category-mode collapse—we propose a non-autoregressive distributional prediction framework. Our approach introduces KL divergence to quantify temporal distributional drift, identifies distributional imbalance as the primary cause of mode collapse, and incorporates three key components: an explicit distribution-aware learning objective, locally order-invariant representations, and multi-token parallel prediction. Experiments across multiple real-world datasets demonstrate substantial improvements over state-of-the-art baselines. The framework offers both interpretability and practicality, providing a principled alternative for behavioral sequence modeling with clear design rationales and deployable mechanisms.
Existing Transformer-based time series forecasting models often suffer from systematic residual biases due to architectural limitations, insufficient modeling of stochasticity, or inadequate multi-scale representations. To address this, this work proposes a two-stage, model-agnostic correction framework: an initial prediction is first generated by a base Transformer, followed by a multi-scale residual-aware meta-corrector that dynamically models structured errors across multivariate channels, enabling end-to-end adaptive refinement. By decoupling forecasting and residual learning into distinct representation stages, the approach formally expands the hypothesis space and overcomes the approximation limitations inherent in single-stage architectures. Evaluated on eight mainstream benchmarks, the method significantly outperforms existing approaches, achieving substantial reductions in both MSE and MAE, effectively mitigating systematic bias and enhancing robustness to complex temporal dynamics.
Existing Transformer-based time series forecasting models overemphasize temporal modeling, resulting in high computational overhead with marginal performance gains and severely constrained representation capacity due to suboptimal embedding design. To address this, we propose Hybrid Temporal–Multivariate Embedding (HTME), a lightweight yet effective embedding framework that jointly captures temporal dynamics and inter-variable dependencies via two synergistic components: a temporal feature extraction module and a multivariate relational modeling module. HTME is plug-and-play compatible with standard Transformer architectures and enhances sequence representation without increasing decoding complexity. Extensive experiments across eight real-world datasets demonstrate that HTME consistently outperforms state-of-the-art baselines—achieving an average 6.2% reduction in MAE and a 31% decrease in FLOPs—thereby validating the fundamental benefit of multivariate-aware embedding for time series modeling.
Existing Transformer-based models for time series forecasting neglect intrinsic temporal properties—namely, unidirectional causality and the decaying influence of past observations over time. To address this, we propose TimeFormer, a novel architecture centered on a Modulated Self-Attention (MoSA) mechanism. MoSA explicitly incorporates temporal priors by integrating a Hawkes process to model decay dynamics, enforcing causal masking to preserve temporal directionality, and enabling multi-scale subsequence analysis to capture dynamic dependencies across granularities. This systematic infusion of temporal inductive bias significantly enhances temporal modeling capability within the Transformer framework. Extensive experiments on multiple real-world benchmarks demonstrate that TimeFormer achieves state-of-the-art performance, outperforming existing methods by up to 7.45% in MSE reduction and setting new records on 94.04% of evaluation metrics. Moreover, the MoSA module exhibits strong transferability, serving as a plug-and-play enhancement for diverse Transformer variants.
This work addresses the limitations of existing foundation time series models, which suffer from high computational overhead, poor adaptability to dynamic data streams, and an inability to learn continuously—hindering their deployment in resource-constrained environments. To overcome these challenges, we propose TimeBlocks, a novel foundation modeling paradigm that uniquely integrates multi-task generalization, lightweight architecture, and continual calibration capabilities. TimeBlocks dynamically assembles compact models at inference time through a modular pool of model blocks and a routing strategy tailored to incoming data streams. Furthermore, it incorporates StreamCore, a streaming summarization algorithm that enables efficient online calibration. Extensive experiments demonstrate that TimeBlocks achieves significantly higher prediction accuracy than current methods across multiple datasets while maintaining low computational costs, enabling effective real-time forecasting under stringent resource constraints.