🤖 AI Summary
Time-series foundation models (TSFMs) and large language model–driven time-series models (TSLLMs) suffer from scarcity of high-quality, diverse real-world time-series data. Method: We systematically investigate the role of synthetic data in pretraining, fine-tuning, and evaluation of TSFMs/TSLLMs. We integrate generative approaches—including GANs, VAEs, diffusion models, LLM-based generation, and prompt-driven synthesis—while explicitly modeling temporal characteristics such as periodicity, abrupt changes, and multi-scale dependencies. Contribution/Results: We propose the first methodology framework for synthetic data in time-series AI, mapping generation strategies to model capability improvements. We categorize seven mainstream synthetic methods and four core application scenarios, identify six critical research gaps, and advocate a future paradigm emphasizing scalability, debiasing, and fidelity. This work delivers the first comprehensive roadmap for synthetic-data–driven time-series AI.
📝 Abstract
Time series analysis is crucial for understanding dynamics of complex systems. Recent advances in foundation models have led to task-agnostic Time Series Foundation Models (TSFMs) and Large Language Model-based Time Series Models (TSLLMs), enabling generalized learning and integrating contextual information. However, their success depends on large, diverse, and high-quality datasets, which are challenging to build due to regulatory, diversity, quality, and quantity constraints. Synthetic data emerge as a viable solution, addressing these challenges by offering scalable, unbiased, and high-quality alternatives. This survey provides a comprehensive review of synthetic data for TSFMs and TSLLMs, analyzing data generation strategies, their role in model pretraining, fine-tuning, and evaluation, and identifying future research directions.