Score
Design, build, and evaluate adaptations of pretrained time-series foundation models by transferring learned sequence representations, fine-tuning or calibrating model components, and applying zero-shot predictions to new series or tasks. This work includes configuring inputs and output heads for forecasting or related time-series outputs, reweighting or adapting pretrained features, and comparing adapted model performance to benchmarks across horizons or assets.
The time-series foundation model (TSFM) field lacks a systematic taxonomy, hindering comparative analysis and principled design. Method: We propose the first multi-dimensional taxonomy tailored to Transformer-based TSFMs, spanning five dimensions: architectural design, forecasting paradigm, variable dimensionality, scale/complexity, and pretraining objective functions—introducing objective function type as a novel classification criterion to unify capability characterization and design rationale. Through comprehensive literature review, architectural analysis, task mapping, and paradigm comparison, we systematically cover mainstream modeling approaches—including patch-based and raw-sequence methods. Contribution/Results: This taxonomy establishes a structured knowledge graph for TSFMs, clarifying technological trends, exposing critical research gaps, and providing a principled foundation for developing scalable, interpretable, and multi-task-cooperative time-series foundation models.
The time-series foundation model (TSM) field suffers from a lack of comprehensive surveys, standardized evaluation protocols, and fragmented technical approaches. Method: This paper systematically reviews over 100 state-of-the-art works published between 2022 and 2024, proposing the first 3E analytical framework for TSMs—emphasizing Effectiveness, Efficiency, and Explainability—and establishing a unified taxonomy spanning application domains, resources, and methodologies. It innovatively categorizes TSM development into two paradigms: “training from scratch” and “large language model (LLM) adaptation,” and introduces a standardized cross-paradigm benchmarking protocol. A multidimensional evaluation suite is designed, integrating structured benchmark datasets, model zoos, and toolchains. Contribution/Results: All components are open-sourced via a GitHub full-stack repository, significantly enhancing reproducibility, comparability, and practical deployment efficiency of TSM research.
To address the weak zero-shot generalization capability in long-horizon time series forecasting, this paper proposes a foundation model framework for time series prediction. Inspired by large language model paradigms, it employs self-supervised pretraining on large-scale heterogeneous time series data to learn general-purpose temporal representations. Crucially, it innovatively unifies point forecasting and probabilistic forecasting within a single modeling objective and introduces a lightweight fine-tuning strategy for downstream adaptation. This approach transcends conventional task-specific architectures, significantly enhancing zero-shot forecasting performance on unseen datasets. Experiments demonstrate that the fine-tuned model achieves an average 18.7% reduction in MAE on long-horizon forecasting tasks. Moreover, it exhibits strong cross-domain adaptability and practical utility across multi-source, multi-frequency, and multi-domain time series. The work establishes a novel paradigm for time series foundation model research.
This work proposes a novel approach to time series forecasting by adapting the large language model paradigm to this domain, introducing a unified foundation model framework built upon large-scale pretraining. Unlike traditional methods that rely on handcrafted architectures with limited generalization, the proposed model supports both point and probabilistic forecasting and is pretrained on diverse time series data. Through carefully designed fine-tuning strategies, it achieves substantial improvements over zero-shot baselines across multiple benchmark datasets, demonstrating the effectiveness and superiority of the pretrain–fine-tune paradigm in time series prediction.
Foundation models (FMs) are increasingly applied to time series forecasting, yet their generalizability across diverse, domain-heterogeneous time series remains poorly understood. Method: This work systematically evaluates FM applicability through cross-domain zero-shot forecasting and fine-tuning experiments, explicitly assessing how pretraining data domain affects out-of-distribution generalization. Contribution/Results: (1) Zero-shot performance is highly sensitive to pretraining domain and degrades sharply on unseen real-world scenarios; (2) After fine-tuning, large FMs fail to consistently outperform lightweight, task-specific models and incur substantially higher computational costs; (3) The widely held “larger models are universally better” assumption is empirically invalidated for time series forecasting, leading to the principle that “domain alignment takes precedence over scale expansion.” These findings provide empirical grounding and methodological guidance for designing practical time-series foundation models—highlighting domain specificity, efficiency trade-offs, and the limits of transferability in sequential forecasting.
This work presents the first systematic investigation of catastrophic forgetting in time-series foundation models (TSFMs) under continual learning settings, where sequential fine-tuning across multiple tasks leads to significant performance degradation on previously learned tasks—revealing a critical robustness deficiency. Method: We propose a knowledge retention quantification framework based on synthetically generated periodic data: we construct controllable periodic synthetic datasets and integrate them with zero-shot transfer and sequential fine-tuning paradigms to decouple and independently assess the stability–plasticity trade-off between new-task adaptation and old-knowledge retention. Contribution/Results: Empirical evaluation demonstrates that while existing TSFMs improve performance on new tasks, they suffer severe forgetting across diverse benchmarks—exposing fundamental limitations in their continual learning capability. Our work establishes a reproducible evaluation benchmark and provides essential diagnostic tools to guide the robust evolution of TSFMs.
To address the challenge of online adaptation to dynamic data distributions after deployment of time-series foundation models, this paper proposes AdapTS, a lightweight online adaptation framework. Methodologically, AdapTS decouples foundation model inference from adapter learning via two novel modules: AdapTS-Forecaster—a linear or MLP-based lightweight time-series modeling component—and AdapTS-Weighter—a gradient-free, dynamically weighted fusion mechanism. It further incorporates online distribution estimation and real-time feedback-driven calibration. Crucially, AdapTS enables zero-shot fine-tuning and low-overhead adaptation without modifying the frozen foundation model. Evaluated across multiple standard benchmarks, AdapTS consistently improves the average MSE of mainstream time-series foundation models by 7.2%–15.8%, while incurring less than 3 ms additional inference latency. This demonstrates substantial gains in both predictive accuracy and practical deployability.
Pretrained time series foundation models often underperform on downstream tasks due to domain shift, task heterogeneity, scarce labeled data, and computational constraints. This work proposes the first systematic post-training framework, categorizing existing approaches along five dimensions based on their intervention points within the forecasting pipeline: parameter adaptation, context augmentation, model composition, output and uncertainty calibration, and compression with specialization. By delineating the design space and inherent limitations of each category, the framework offers a structured pathway to bridge the gap between pretraining and reliable deployment, thereby advancing the standardization and systematic development of time series post-training methodologies.
Existing time series pretraining methods struggle to generalize effectively across multiple datasets due to discrepancies in input length and channel dimensions. This work proposes ADAPT, a novel pretraining paradigm that enables unified modeling across 162 time series classification datasets by adaptively aligning the physical attributes of time series data. Integrating self-supervised learning with a hybrid batch training strategy, ADAPT overcomes the generalization limitations inherent in conventional many-to-one pretraining approaches. The method achieves state-of-the-art performance on multiple benchmarks, establishing a foundational framework for developing general-purpose foundation models for time series analysis.
This work addresses the challenge of online adaptation for black-box time series foundation models when access to internal parameters is unavailable. The authors propose ORCA, a novel approach that explicitly learns the mapping between prediction errors and input-output contextual information, enabling dynamic correction of forecasts through residual modeling—without modifying the original model or relying on gradient-based updates. By circumventing the conventional paradigm of white-box fine-tuning, ORCA demonstrates superior performance over existing black-box adaptation strategies across five state-of-the-art foundation models and eight benchmark datasets, establishing its effectiveness and broad applicability.
This work addresses the performance degradation of time series foundation models (TSFMs) in zero-shot forecasting due to domain shifts. To mitigate this issue, the authors propose MixFT, a novel approach that leverages a Bayesian mixture model to partition source data into more homogeneous latent subdomains. For each inferred subdomain, a lightweight LoRA module is independently fine-tuned, enabling fine-grained adaptation of the TSFM. Unlike conventional strategies that apply uniform fine-tuning across the entire dataset or rely on predefined dataset splits, MixFT more accurately captures the distinct characteristics of individual subdomains. Extensive experiments demonstrate that MixFT significantly outperforms existing methods across multiple zero-shot forecasting tasks, thereby validating the efficacy and novelty of subdomain-specialized fine-tuning.
This work addresses the pervasive structural redundancy in time series foundation models (TSFMs), which undermines their reliability and efficiency. Through large-scale evaluation and mechanistic interpretability analysis, we find that mainstream TSFMs exhibit robustness to entire-layer removal and identify specific attention heads responsible for dominant repetitive patterns and seasonal biases. We propose an intrinsic pruning strategy based on stable rank, offering the first systematic characterization of shared redundancy mechanisms and their degradation origins across diverse TSFMs. By integrating component ablation, direct logit attribution via residual stream analysis, and a theoretical framework interpreting Transformers as kernel regressors, we validate both the layer-wise redundancy and the critical role of particular attention heads across multiple real-world and synthetic datasets, thereby establishing a new pathway toward efficient and reliable time series modeling.