Score
Designs and trains self-supervised or unsupervised pretraining pipelines and model architectures for multivariate time-series data that learn temporal and cross-variable representations using masked reconstruction or related masking-based objectives, producing foundation models or frozen encoders. These artifacts and procedures support large-scale pretraining and subsequent transfer learning, zero-shot use, or fine-tuning on downstream time-series tasks.
The time-series foundation model (TSFM) field lacks a systematic taxonomy, hindering comparative analysis and principled design. Method: We propose the first multi-dimensional taxonomy tailored to Transformer-based TSFMs, spanning five dimensions: architectural design, forecasting paradigm, variable dimensionality, scale/complexity, and pretraining objective functions—introducing objective function type as a novel classification criterion to unify capability characterization and design rationale. Through comprehensive literature review, architectural analysis, task mapping, and paradigm comparison, we systematically cover mainstream modeling approaches—including patch-based and raw-sequence methods. Contribution/Results: This taxonomy establishes a structured knowledge graph for TSFMs, clarifying technological trends, exposing critical research gaps, and providing a principled foundation for developing scalable, interpretable, and multi-task-cooperative time-series foundation models.
To address feature extraction, anomaly detection, and cross-domain classification challenges posed by large-scale unlabeled time-series data in wireless communications, radar, biomedical applications, and IoT, this paper presents a systematic survey and proposes a novel hybrid modeling and self-supervised co-optimization framework tailored for time-series signals. It innovatively integrates convolutional, recurrent, and temporal convolutional autoencoders with vision Transformers (e.g., ViT, TimeSformer), incorporating masked time-series modeling, contrastive learning, and multi-scale feature fusion to enhance model interpretability and domain generalizability. Evaluated on a unified benchmark across ECG, radar waveforms, and IoT sensor datasets, the framework achieves an average 12.6% improvement in anomaly detection F1-score and a 9.3% gain in classification accuracy over state-of-the-art baselines.
Existing time series pretraining methods struggle to generalize effectively across multiple datasets due to discrepancies in input length and channel dimensions. This work proposes ADAPT, a novel pretraining paradigm that enables unified modeling across 162 time series classification datasets by adaptively aligning the physical attributes of time series data. Integrating self-supervised learning with a hybrid batch training strategy, ADAPT overcomes the generalization limitations inherent in conventional many-to-one pretraining approaches. The method achieves state-of-the-art performance on multiple benchmarks, establishing a foundational framework for developing general-purpose foundation models for time series analysis.
This work explores the potential of non-contrastive learning for pretraining foundation models on time series data to enhance downstream classification performance. We introduce, for the first time, a DINOv2-style self-distillation mechanism into time series modeling, integrating the Mantis tokenizer with a Transformer encoder within a student-teacher framework. This approach jointly optimizes temporal invariance and local fine-grained structure through temporal cropping augmentation and block-wise masked reconstruction, enabling multi-objective pretraining. Extensive experiments on the UCR and UEA benchmarks demonstrate that our method achieves state-of-the-art performance, validating the effectiveness and superiority of non-contrastive self-distillation for time series representation learning.
This study addresses the lack of systematic, quantitative evaluation of the practical benefits conferred by self-supervised pretraining on time series across diverse downstream tasks. The authors construct a controlled evaluation framework to systematically compare generative and latent alignment–based approaches, introducing a novel data augmentation strategy based on the discrete wavelet transform (DWT) to enhance invariance to local perturbations. They provide the first quantification of the asymmetry in “pretraining gains,” revealing that representation utility is governed by a trade-off between signal resolution required by the task and the precision–invariance balance inherent in the learning objective. Representation quality is found to be independent of data provenance and saturates at moderate model depths. Experiments demonstrate pretraining improvements of up to 375% on anomaly detection and classification tasks, yet limited gains in forecasting, while also validating the efficacy of scaling models with large-scale synthetic data.
To address the challenges of complex feature fusion and poor scalability in unsupervised multimodal time-series representation learning, this paper proposes a spectral-graph-theoretic framework for implicit cross-modal alignment. Methodologically, it abandons explicit multimodal feature fusion and instead adopts a single-encoder, multi-view architecture: diverse time-series views—such as frequency-domain, image-based, and symbolic representations—are generated via modality-agnostic temporal transformations, and spectral-graph-guided contrastive loss enables unsupervised cross-modal alignment. This lightweight design implicitly captures latent inter-modal dependencies, strengthening inductive bias and generalization capacity. Extensive experiments across multiple domains demonstrate state-of-the-art performance on downstream tasks—including classification, forecasting, and anomaly detection—surpassing existing unsupervised methods by an average accuracy gain of 3.2%–7.8%.
Existing time-series self-supervised learning methods struggle to simultaneously model long-term dynamics and capture fine-grained local patterns. To address this, we propose an autoregressive pretraining framework that integrates causal Transformers with denoising diffusion: causal Transformers enable long-range dependency modeling, while block-wise embedding and an autoregressive diffusion process—comprising forward noise injection and reverse denoising—jointly optimize global trend and local structural representations. This work is the first to unify global causal modeling and local fine-grained reconstruction within a single self-supervised objective, enhancing representation discriminability and robustness through end-to-end joint training. Extensive experiments on multiple time-series forecasting and classification benchmarks demonstrate significant improvements over state-of-the-art methods, validating the framework’s strong transferability and generalization capability.
This study addresses the bottleneck in time-series self-supervised pretraining where high-frequency noise disrupts structural learning and visual augmentation techniques are difficult to transfer directly. To overcome this, we propose a novel time-frequency self-distillation framework tailored for temporal signals. Centered on the wavelet transform, our method constructs invariant views through multi-scale time-frequency augmentations rather than conventional spatial transformations, enabling architecture-agnostic and efficient representation learning. Extensive experiments demonstrate that the proposed model surpasses state-of-the-art methods across long-term forecasting, zero-shot transfer, and anomaly detection tasks. Notably, linear probing on frozen representations outperforms fully supervised baselines, validating both the effectiveness and generalization capability of the proposed paradigm.
This work investigates the generalization performance and spectral structure of matrix-valued predictors formed by aggregating multiple masks in masked self-supervised learning under high-dimensional settings. Leveraging random matrix theory within an asymptotic framework where sample size and dimension grow proportionally, the study establishes the first high-dimensional theoretical analysis for masked self-supervised learning. The core contributions include deriving an explicit expression for the generalization error, characterizing the spectral properties of the aggregated predictor, revealing a BBP-type phase transition under spiked covariance models, and identifying the precise threshold conditions under which latent signals can be effectively recovered. The analysis further provides theoretical evidence that this approach outperforms classical PCA in certain structured scenarios.
Clinical time series analysis faces significant challenges including limited sample sizes, data heterogeneity, and protocol drift, necessitating generalizable representation learning approaches that jointly support classification and regression tasks. This work proposes PathoFM, a framework that systematically investigates the impact of inductive biases on representation transfer using gait data from spinal cord injury patients. The approach introduces a multi-objective self-supervised pretraining strategy that integrates local structural reconstruction, causal temporal continuity, and individual-specific contextual conditioning, all jointly optimized within a Transformer encoder. Experimental results demonstrate that this dynamics-driven hybrid objective substantially outperforms single-objective methods, achieving superior generalization performance across both cross-task and cross-subject settings.
本文通过引入新的理论框架分析了掩码预训练(MPT)的工作机制,提出了均匀性增强的MPT损失(U-MPT)以解决维度坍缩问题,并提出了一种新的掩码策略来提高下游任务性能。
This study investigates whether self-supervised pre-training (SPT) can effectively enhance the diagnostic performance of Transformer models on multimodal, multivariate, and univariate medical time-series data, particularly in data-scarce clinical settings. The work proposes a general-purpose approach that requires no task-specific architectural modifications and incorporates four masking strategies to facilitate representation learning across both modalities and temporal dimensions. Experimental results across three medical time-series tasks demonstrate that SPT improves classification accuracy by 0–6 percentage points, with deeper models exhibiting more pronounced gains. Notably, this is the first study to validate the efficacy of SPT on univariate medical time-series data, demonstrating its strong generalizability, scalability, and robustness under limited-data conditions.