Score
Designs and builds forecasting models and operational pipelines that produce multi-horizon predictions for new, unseen time series without per-series retraining by adapting time-series foundation models or LLMs; these systems handle hierarchical structures and irregular sampling, emit forecasts for multiple series/keys at once, and provide concise forecast summaries for downstream decision contexts. Also develops evaluation and diagnostic tooling (e.g., rolling-origin evaluation) to analyze and validate generalization across series, hierarchy levels, and forecast horizons.
The time-series foundation model (TSM) field suffers from a lack of comprehensive surveys, standardized evaluation protocols, and fragmented technical approaches. Method: This paper systematically reviews over 100 state-of-the-art works published between 2022 and 2024, proposing the first 3E analytical framework for TSMs—emphasizing Effectiveness, Efficiency, and Explainability—and establishing a unified taxonomy spanning application domains, resources, and methodologies. It innovatively categorizes TSM development into two paradigms: “training from scratch” and “large language model (LLM) adaptation,” and introduces a standardized cross-paradigm benchmarking protocol. A multidimensional evaluation suite is designed, integrating structured benchmark datasets, model zoos, and toolchains. Contribution/Results: All components are open-sourced via a GitHub full-stack repository, significantly enhancing reproducibility, comparability, and practical deployment efficiency of TSM research.
The time-series foundation model (TSFM) field lacks a systematic taxonomy, hindering comparative analysis and principled design. Method: We propose the first multi-dimensional taxonomy tailored to Transformer-based TSFMs, spanning five dimensions: architectural design, forecasting paradigm, variable dimensionality, scale/complexity, and pretraining objective functions—introducing objective function type as a novel classification criterion to unify capability characterization and design rationale. Through comprehensive literature review, architectural analysis, task mapping, and paradigm comparison, we systematically cover mainstream modeling approaches—including patch-based and raw-sequence methods. Contribution/Results: This taxonomy establishes a structured knowledge graph for TSFMs, clarifying technological trends, exposing critical research gaps, and providing a principled foundation for developing scalable, interpretable, and multi-task-cooperative time-series foundation models.
This study evaluates whether foundation models can replace traditional supervised methods for real-world time series forecasting without task-specific training. It introduces a novel operational perspective by categorizing forecasting tasks into four representative scenarios and proposes a sequence-feature-based complexity-aware routing mechanism to automatically select the optimal model. Through extensive cross-domain benchmarking and zero-shot inference comparisons against supervised baselines, the work demonstrates that foundation models excel in settings with transferable periodic structures or cold-start conditions. The proposed routing strategy not only maintains competitive prediction accuracy but also significantly reduces inference overhead, outperforming uniform deployment of foundation models across all tasks.
To address the challenge of online adaptation to dynamic data distributions after deployment of time-series foundation models, this paper proposes AdapTS, a lightweight online adaptation framework. Methodologically, AdapTS decouples foundation model inference from adapter learning via two novel modules: AdapTS-Forecaster—a linear or MLP-based lightweight time-series modeling component—and AdapTS-Weighter—a gradient-free, dynamically weighted fusion mechanism. It further incorporates online distribution estimation and real-time feedback-driven calibration. Crucially, AdapTS enables zero-shot fine-tuning and low-overhead adaptation without modifying the frozen foundation model. Evaluated across multiple standard benchmarks, AdapTS consistently improves the average MSE of mainstream time-series foundation models by 7.2%–15.8%, while incurring less than 3 ms additional inference latency. This demonstrates substantial gains in both predictive accuracy and practical deployability.
This study addresses ongoing skepticism regarding the effectiveness of large language models (LLMs) in time series forecasting. Through a large-scale empirical evaluation encompassing 8 billion observations across 17 diverse scenarios, the authors propose LLM4TS—a framework integrating pre-alignment and post-alignment strategies, token-level routing analysis, and prompt engineering. Their findings demonstrate that pre-alignment consistently outperforms post-alignment in over 90% of tasks, revealing complementary roles between LLMs’ pre-trained knowledge and architectural design. Moreover, full-scale LLMs significantly enhance performance in mixed-distribution and cross-domain generalization settings, thereby affirming their critical value in time series prediction.
This study systematically investigates the capabilities and underlying mechanisms of large language models (LLMs) in zero-shot time series forecasting. It addresses their observed bias toward periodic/trended sequences and sharp performance degradation on irregular, non-stationary data. To tackle these limitations, the authors propose two novel methodological contributions: (1) the first empirical identification of LLMs’ implicit periodicity detection capability, coupled with a cross-modal time-series–text representation analysis framework; and (2) a knowledge-enhanced zero-shot prompting paradigm integrating natural-language rephrasing and external domain knowledge injection. Experiments demonstrate that while LLMs achieve accuracy comparable to classical statistical and deep learning baselines on strongly periodic series, their performance deteriorates significantly on non-periodic data. With knowledge augmentation, average forecasting error decreases by 23.6%. This work advances the understanding of LLMs’ temporal reasoning capacity and establishes a reproducible, interpretable prompting framework for time series forecasting.
To address modeling and evaluation challenges posed by heterogeneous data—characterized by multi-frequency sampling, high dimensionality, and multimodality—in large time-series models (LTSMs), this paper introduces the first unified toolbox and benchmark platform for time-series forecasting. Methodologically, it achieves full-stack decoupling and co-evaluation across preprocessing, tokenization, prompt learning, training paradigms, and data diversity; it further proposes a Transformer-based autoregressive architecture, multi-granularity tokenization, instruction-style prompting, and cross-frequency/dimension adaptation techniques. The core contribution lies in systematically uncovering strong synergistic effects among design choices and identifying an optimal configuration, which significantly improves zero-shot and few-shot generalization performance across multiple standard benchmarks—outperforming both state-of-the-art LTSMs and conventional time-series models.
研究基于大语言模型的预测代理,通过结合语言推理、时间数据等方法解决未来或未观测目标的预测问题,并探讨了训练、评估及应用。
This work addresses the “last-mile” gap between statistical time series forecasting and business decision-making—specifically, the challenge of effectively incorporating weakly structured contextual factors such as holidays and marketing campaigns. To bridge this gap, we propose the first framework that systematically integrates large language model (LLM) agents into the post-forecasting refinement stage. Our approach unifies the forecasting workspace, leverages tool-augmented retrieval of external evidence, and translates LLM reasoning into explicit revision operations under structured safety constraints. It further supports Map-Reduce-style long-horizon divide-and-conquer forecasting and incorporates a memory-based reflection mechanism. Evaluated in real-world business settings, the method significantly enhances the operational relevance and decision utility of forecasts, delivering controllable, auditable, and business-ready predictions.
This work proposes a training-free framework that formulates time series forecasting as a planning problem, synergistically integrating the textual reasoning capabilities of large language models (LLMs) with the numerical prediction power of frozen time series foundation models (TSFMs), such as Chronos or TimesFM. The approach employs a TSFM as a trajectory simulator to generate candidate forecasts, while two role-specialized LLMs act as a policy (Ranker) and a value function (Judge), respectively. Guided by Monte Carlo Tree Search (MCTS), the method selects the optimal forecast trajectory under natural language conditions while preserving temporal structure. Experiments on the Context-is-Key and Time-MMD benchmarks demonstrate consistent and significant performance gains across diverse TSFM–LLM pairings, establishing the first training-free, cross-modal framework for text-conditioned time series forecasting.
Pretrained time series foundation models often underperform on downstream tasks due to domain shift, task heterogeneity, scarce labeled data, and computational constraints. This work proposes the first systematic post-training framework, categorizing existing approaches along five dimensions based on their intervention points within the forecasting pipeline: parameter adaptation, context augmentation, model composition, output and uncertainty calibration, and compression with specialization. By delineating the design space and inherent limitations of each category, the framework offers a structured pathway to bridge the gap between pretraining and reliable deployment, thereby advancing the standardization and systematic development of time series post-training methodologies.
This work proposes a novel approach to time series forecasting by adapting the large language model paradigm to this domain, introducing a unified foundation model framework built upon large-scale pretraining. Unlike traditional methods that rely on handcrafted architectures with limited generalization, the proposed model supports both point and probabilistic forecasting and is pretrained on diverse time series data. Through carefully designed fine-tuning strategies, it achieves substantial improvements over zero-shot baselines across multiple benchmark datasets, demonstrating the effectiveness and superiority of the pretrain–fine-tune paradigm in time series prediction.