🤖 AI Summary
This study addresses the absence of effective evaluation mechanisms for learning signals during data sampling in time series foundation model pretraining. To this end, it proposes a static data selection framework that employs a reference predictor to score samples and retains those within an intermediate interval. By establishing a theoretical connection between loss and gradient norms, and integrating stratified dataset selection with local Jacobian conditioning analysis, the approach achieves efficient data filtering across varying scales and architectures. This method substantially reduces the number of pretraining windows required while significantly improving MASE and CRPS metrics, outperforming baselines trained on full datasets. Consequently, it establishes an efficient new paradigm for data selection in time series foundation models.
📝 Abstract
Time series foundation models (TSFMs) are pretrained on heterogeneous collections containing billions of observations, yet their training windows are typically sampled without estimating whether they provide useful learning signal. We introduce a static data-selection framework that scores each window with a reference forecaster and retains an intermediate interval within every source dataset. Specifically, we connect forecasting loss to optimization difficulty by showing that normalized squared loss controls the per-sample gradient norm under a local Jacobian condition. We then define a reference loss score and apply dataset-stratified selection to preserve the diversity of samples. Across various TSFM architectures, our method outperforms random selection by an absolute margin and even improves both relative MASE and CRPS over full-data pretraining by retaining fewer candidate pretraining windows. Further analyses show strong cross-scale and cross-architecture score correlations, indicating that a small reference model can often select data for larger targets, provided that the reference and target share compatible difficulty orderings.