Empowering Time Series Analysis with Synthetic Data: A Survey and Outlook in the Era of Foundation Models

📅 2025-03-14
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Time-series foundation models (TSFMs) and large language model–driven time-series models (TSLLMs) suffer from scarcity of high-quality, diverse real-world time-series data. Method: We systematically investigate the role of synthetic data in pretraining, fine-tuning, and evaluation of TSFMs/TSLLMs. We integrate generative approaches—including GANs, VAEs, diffusion models, LLM-based generation, and prompt-driven synthesis—while explicitly modeling temporal characteristics such as periodicity, abrupt changes, and multi-scale dependencies. Contribution/Results: We propose the first methodology framework for synthetic data in time-series AI, mapping generation strategies to model capability improvements. We categorize seven mainstream synthetic methods and four core application scenarios, identify six critical research gaps, and advocate a future paradigm emphasizing scalability, debiasing, and fidelity. This work delivers the first comprehensive roadmap for synthetic-data–driven time-series AI.

Technology Category

Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageMachine Learning: Time-Series/Data StreamsPlanning, Routing, and Scheduling: Temporal Planning

Application Category

Social Networks and Social Media: Generative AI / large language models and their impact on social systemsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsWeb Mining and Content Analysis: Web data generation and simulation
📝 Abstract
Time series analysis is crucial for understanding dynamics of complex systems. Recent advances in foundation models have led to task-agnostic Time Series Foundation Models (TSFMs) and Large Language Model-based Time Series Models (TSLLMs), enabling generalized learning and integrating contextual information. However, their success depends on large, diverse, and high-quality datasets, which are challenging to build due to regulatory, diversity, quality, and quantity constraints. Synthetic data emerge as a viable solution, addressing these challenges by offering scalable, unbiased, and high-quality alternatives. This survey provides a comprehensive review of synthetic data for TSFMs and TSLLMs, analyzing data generation strategies, their role in model pretraining, fine-tuning, and evaluation, and identifying future research directions.
Problem

Research questions and friction points this paper is trying to address.

Addressing dataset challenges for time series analysis models
Exploring synthetic data for scalable and unbiased alternatives
Reviewing synthetic data's role in model training and evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Synthetic data enhances time series analysis.
Foundation models enable generalized learning.
Data generation strategies improve model training.
🔎 Similar Papers
No similar papers found.