Score
Designs and implements methods and pipelines to generate, resample, and transform synthetic time-series data using stochastic and deterministic models, interpolation/extrapolation techniques, and bootstrapping so as to produce benchmark datasets for backtesting and comparison. Builds evaluation and validation procedures — metrics, statistical tests, and simulation-based experiments — to assess time-series modeling, forecasting, and regression methods under controlled, repeatable conditions.
Time-series foundation models (TSFMs) and large language model–driven time-series models (TSLLMs) suffer from scarcity of high-quality, diverse real-world time-series data. Method: We systematically investigate the role of synthetic data in pretraining, fine-tuning, and evaluation of TSFMs/TSLLMs. We integrate generative approaches—including GANs, VAEs, diffusion models, LLM-based generation, and prompt-driven synthesis—while explicitly modeling temporal characteristics such as periodicity, abrupt changes, and multi-scale dependencies. Contribution/Results: We propose the first methodology framework for synthetic data in time-series AI, mapping generation strategies to model capability improvements. We categorize seven mainstream synthetic methods and four core application scenarios, identify six critical research gaps, and advocate a future paradigm emphasizing scalability, debiasing, and fidelity. This work delivers the first comprehensive roadmap for synthetic-data–driven time-series AI.
This study addresses a key limitation of traditional time series scenario generation methods, which often produce interest rate curve paths lacking economic plausibility despite matching historical return distributions. To overcome this, the authors propose a novel framework that integrates parametric term structure models with a semi-parametric bootstrap approach. The method preserves the dynamic structure of the underlying model while resampling residuals and incorporates autoregressive or mean-reverting specifications to enhance temporal coherence of generated paths. Empirical results demonstrate that, in fixed-income applications, the proposed approach significantly outperforms conventional nonparametric techniques by simultaneously reproducing the statistical properties of yield curves accurately and generating scenarios that align with economic reasoning.
To address the lack of objective, large-scale benchmarks for evaluating synthetic time series, this paper introduces STEB—the first standardized evaluation benchmark. STEB comprises ten real and synthetic datasets, integrates thirteen configurable transformations and stochasticity-injection mechanisms, and supports parallel/sequential evaluation alongside runtime and error tracking. We propose a dual-dimension evaluation framework—“reliability” and “consistency”—and systematically benchmark 41 mainstream evaluation metrics for the first time. Empirical analysis reveals that time-series embeddings exert a decisive influence on metric outcomes: different embeddings substantially alter metric rankings and score stability. The open-source STEB framework includes a fully reproducible evaluation pipeline, establishing a scalable, interpretable, and automated paradigm for synthetic time-series assessment.
To address insufficient training data and distribution imbalance in few-shot multivariate time series forecasting, this paper proposes the first online data augmentation framework embedded within the training iteration process: synthetic samples are generated dynamically and paired with real samples within each mini-batch, thereby avoiding distributional shift and storage overhead associated with offline augmentation. The method unifies three backbone architectures—TCN, LSTM, and Informer—and integrates seven synthesis techniques, including TS-TCC, GAN-based generation, and diffusion-inspired approaches, augmented by an adaptive online sampling scheduling strategy. Evaluated on six benchmark datasets comprising 3,797 time series, the framework achieves an 8.2% reduction in MASE and a 6.7% reduction in MAE compared to both no-augmentation and offline-augmentation baselines, demonstrating significant improvements in prediction accuracy, generalization, and robustness under limited-data regimes.
Pretrained time series foundation models often underperform on downstream tasks due to domain shift, task heterogeneity, scarce labeled data, and computational constraints. This work proposes the first systematic post-training framework, categorizing existing approaches along five dimensions based on their intervention points within the forecasting pipeline: parameter adaptation, context augmentation, model composition, output and uncertainty calibration, and compression with specialization. By delineating the design space and inherent limitations of each category, the framework offers a structured pathway to bridge the gap between pretraining and reliable deployment, thereby advancing the standardization and systematic development of time series post-training methodologies.
Traditional bootstrap and conformal prediction methods fail in time series settings due to violations of exchangeability and the absence of a unified framework that supports dependence-aware resampling and adaptive conformal calibration. This work proposes the first typed API integrating block, residual, sieve, and wild resampling schemes with adaptive conformal approaches such as EnbPI and ACI, enabling distribution-free uncertainty quantification. Leveraging compilation-based acceleration and streaming reductions, the method requires only O(B) additional memory, circumventing the O(Bn) tensor duplication typical of conventional implementations. Empirical results demonstrate that the approach substantially mitigates undercoverage under the i.i.d. assumption, with sieve resampling achieving coverage closest to the nominal level for short-memory linear processes, while running several times faster than the arch benchmark.
This study addresses the challenge of uncertainty quantification in aggregated time series forecasting, particularly for annual totals and year-over-year growth rates. It proposes a simulation-augmented multi-step split conformal prediction method (SA-MSCP), which generates future trajectories via block bootstrap resampling from cross-validated residuals and constructs calibrated prediction intervals using empirical quantiles. By innovatively integrating a simulation-augmentation mechanism into the multi-step split conformal prediction framework, the method significantly improves empirical coverage for both aggregate totals and their growth rates, yielding more reliable uncertainty estimates without compromising predictive accuracy.
Existing approaches struggle to effectively quantify the similarity and quality between synthetic and real data in evaluating tool-augmented agents. To address this gap, this work proposes SynAE, a novel framework that establishes the first multi-axis evaluation system tailored for multi-turn tool-use scenarios. SynAE introduces four fine-grained metric categories—assessing task instructions, tool invocations, final outputs, and downstream evaluation performance—to systematically measure synthetic data across dimensions of validity, fidelity, and diversity. Integrating natural language processing, trajectory modeling, and controllable generation techniques, the framework enables a reproducible evaluation pipeline and successfully identifies several representative failure modes in synthetic data generation. Empirical results demonstrate that such multidimensional assessment is essential for enhancing the reliability of agent evaluations.
This study addresses the lack of systematic evaluation of stationarity-inducing transformations across diverse non-stationary time series. The authors construct synthetic datasets encompassing trend, seasonality, and heteroskedasticity, complemented by real-world airport passenger flow data, and conduct 3,528 controlled experiments evaluating 14 transformation methods across seven forecasting models and three prediction horizons. Innovatively, stationarity is assessed via consensus from ten statistical tests, and mediation analysis elucidates underlying mechanisms. Results challenge the common assumption that transformations universally improve forecasts: matched transformations enhance accuracy in only 18% of cases; log or Box–Cox transformations are effective for heteroskedastic data (60–65% of cases); and differencing consistently degrades performance on linear-trend series.