Score
Design and implement encoders and representation-learning pipelines that extract and compress multi-timescale temporal and temporal–spectral features from sequential data, using constructs such as multi-scale TCNs, time–frequency encoders, and temporal contrastive objectives. Analyze and optimize these models for cross-scale alignment, domain-transferability, computational and compression efficiency, and redundancy reduction to produce compact, transferable temporal representations.
Deep learning for time series lacks standardized evaluation benchmarks and systematic architectural analysis. Method: We introduce TSLib, the first standardized deep time-series benchmark library, encompassing 24 state-of-the-art models, 30 cross-domain datasets, and five core tasks. It decouples modeling paradigms along two dimensions—fundamental modules and holistic architectures—to establish a modular, reproducible, and fair evaluation framework. Contribution/Results: Through extensive experiments, we uncover, for the first time, strong structural-task alignment patterns—i.e., specific architectural characteristics (e.g., attention mechanisms, convolutional depth, or recurrence design) exhibit consistent performance advantages on particular task types (e.g., forecasting, classification, anomaly detection). This finding provides empirical guidance for model selection and architecture design. Comprehensive evaluation across 12 advanced models validates the robustness of this insight. All code, data, and evaluation pipelines are publicly released and have garnered significant attention from the research community.
This study systematically evaluates the practical performance of time-series embedding methods for classification tasks, providing empirical guidance for method selection. We conduct a unified benchmark across diverse real-world datasets, assessing five embedding paradigms—shape-based, statistical, spectral, deep learning–based, and graph-structured—paired with downstream classifiers including SVM, random forests, CNNs, and LSTMs. To our knowledge, this is the first reproducible, cross-paradigm comparison of embedding techniques within a standardized classification pipeline, revealing that embedding efficacy critically depends on both data characteristics and classifier compatibility. We release a lightweight, modular open-source framework enabling rapid validation and industrial deployment of embedding–classification pipelines. Key contributions include: (1) the first taxonomy of time-series embedding methods specifically designed for classification; (2) the most comprehensive empirical benchmark to date; and (3) publicly available, well-documented code to facilitate reproducibility and application-specific adaptation.
Time-series modeling faces challenges including variable-length sequence handling, high feature redundancy, and limited generalization capability. To address these, we propose ConvFormer—a convolution-like multi-scale fusion framework that jointly employs temporal patching and multi-head attention to progressively compress the time dimension while expanding channel capacity. It introduces cross-scale attention and logarithmic-space normalization to enhance multi-scale feature interaction and suppress redundant representations. The resulting hierarchical time-series representation achieves significant improvements over state-of-the-art Transformer- and CNN-based baselines on both forecasting and classification tasks, reducing feature redundancy by 12.6%–28.4% and improving average performance by 3.7%–9.2%. Our core contribution lies in the first unified integration of convolution-like structural inductive bias, cross-scale attention, and logarithmic normalization within a Transformer architecture—effectively balancing local pattern modeling with global dependency capture.
Existing video classification methods naively average frame-level embeddings from pretrained Transformer encoders, neglecting critical temporal structures—including event ordering, dynamic feature importance, and duration variability. While mainstream temporal modeling approaches require architectural modifications and full retraining, they are incompatible with already fine-tuned large models. This paper proposes an encoder-agnostic, lightweight temporal matching framework: it maps fixed-length embeddings into variable-length multivariate time series and introduces a learnable per-frame, per-feature weighting mechanism. Inspired by time-series alignment, the framework employs a dedicated neural architecture for temporal modeling. It adds fewer than 1.8% parameters and requires no encoder modification or retraining. Evaluated on Something-Something V2, Kinetics-400, and HMDB51, our method achieves 77.2%, 89.1%, and 88.6% Top-1 accuracy, respectively, with training completed in under three hours.
Real-world multivariate time series exhibit strong non-stationarity and cross-scale dynamics; however, prevailing models rely on fixed-scale priors—such as chunk-based tokenization or static frequency-domain transformations—limiting modeling flexibility and hindering robustness to abrupt, high-magnitude events. To address this, we propose an adaptive hierarchical architecture that jointly captures instantaneous fluctuations and long-term trends via multi-scale convolutional encoding, integrates sequential modeling using BiLSTM or Transformer backbones, incorporates Squeeze-and-Excitation gating for channel-wise feature recalibration, and employs multi-head temporal attention for context-aware dynamic feature fusion. The resulting framework establishes a unified paradigm for time-series modeling. Evaluated across 32 benchmark datasets on forecasting, imputation, and classification tasks, our method achieves state-of-the-art performance on 24 datasets—outperforming leading approaches including EMTSF, TimesNet, and PatchTST by significant margins.
To address the challenges of complex feature fusion and poor scalability in unsupervised multimodal time-series representation learning, this paper proposes a spectral-graph-theoretic framework for implicit cross-modal alignment. Methodologically, it abandons explicit multimodal feature fusion and instead adopts a single-encoder, multi-view architecture: diverse time-series views—such as frequency-domain, image-based, and symbolic representations—are generated via modality-agnostic temporal transformations, and spectral-graph-guided contrastive loss enables unsupervised cross-modal alignment. This lightweight design implicitly captures latent inter-modal dependencies, strengthening inductive bias and generalization capacity. Extensive experiments across multiple domains demonstrate state-of-the-art performance on downstream tasks—including classification, forecasting, and anomaly detection—surpassing existing unsupervised methods by an average accuracy gain of 3.2%–7.8%.
This work addresses the limitations of univariate time series classification stemming from singular input representations and insufficient model scalability by proposing LiteMV, the first scalable framework that jointly leverages multi-representation learning and multi-scale modeling. The framework comprises three core components: MSNet, which emphasizes robustness and probabilistic calibration; LS-Net, an efficient and lightweight architecture; and a LiteMV cross-representation interaction mechanism tailored for univariate signals. Through systematic integration of multi-scale convolutions, multi-representation fusion, and calibration-aware optimization, extensive evaluation across 142 benchmark datasets demonstrates that LiteMV achieves the highest average accuracy, MSNet attains the best calibration performance (lowest negative log-likelihood), and LS-Net offers an optimal trade-off between accuracy and efficiency. Pareto analysis further confirms the framework’s ability to flexibly balance accuracy, calibration quality, and computational constraints.
Scientific time series data are often sparse, heterogeneous, and limited in scale, posing significant challenges for unified representation learning. To address this, this work proposes a cross-domain knowledge distillation framework that systematically integrates complementary knowledge from foundation models across multiple scientific domains to construct a universal encoder. The approach introduces an adaptive chunking strategy to handle variable sequence lengths and incorporates a statistical compensation mechanism to mitigate discrepancies arising from numerical scale variations. Extensive experiments across seven scientific time series tasks demonstrate that the proposed method substantially enhances model generalization and transferability, establishing a new paradigm for representation learning in scientific data.
Standard attention mechanisms, constrained by their convex combination formulation, struggle to model signed and oscillatory global temporal operators—such as filtering and harmonic structures—in time series, limiting their performance. This work proposes the Temporal Operator Attention (TOA) framework, which transcends the simplex constraint imposed by softmax by introducing learnable, explicit sequence-space operators. TOA enables input-adaptive, signed cross-timestep mixing and incorporates stochastic operator regularization to enhance training stability. The framework seamlessly integrates into backbone architectures like PatchTST and iTransformer, consistently outperforming baseline methods across forecasting, anomaly detection, and classification tasks, with particularly pronounced gains in reconstruction-intensive scenarios.
Existing neural processes often suffer from underfitting and limited generalization when modeling strongly periodic or quasi-periodic time series, spatial, and image data. This work proposes a Spectral-Aware Transformer Neural Process that enhances periodicity modeling by estimating the spectral characteristics of context points via a spectral aggregator and compressing them into a task-adaptive spectral mixture model, which is then fused with time-domain embeddings. Innovatively, a spectral mixture kernel prior is incorporated into the neural process to reshape the geometry of latent-space similarity, ensuring that points distant in Euclidean space but close in periodic structure remain proximate—thereby overcoming the limitations of conventional translation-equivariant assumptions. Experiments demonstrate that the proposed method significantly outperforms current baselines on both synthetic and real-world time series and image datasets, achieving superior predictive performance for periodic data.
This work addresses the challenges of feature shift, temporal drift, and spectral discrepancy in source-free domain adaptation for time series. To tackle these issues without access to source-domain data, the authors propose a time–frequency joint alignment approach that explicitly corrects spectral shifts—a first in the source-free setting. By modeling the source domain’s temporal dependencies and spectral characteristics at multiple scales, they design a trainable frequency-domain adaptation module that modulates both phase and amplitude of target-domain signals to achieve distribution alignment. Extensive experiments demonstrate that the proposed method significantly outperforms existing source-free time series adaptation techniques across multiple benchmark datasets, confirming its effectiveness and robustness.