🤖 AI Summary
This study addresses the parameter explosion in time-series foundation models caused by multi-scale representations by proposing a fractal weight-sharing mechanism grounded in temporal self-similarity. The method reuses operators across scales via geometrically expanding ladders and integrates parameter-free seasonality detection, shared local patch encoding, and an explicit seasonal decoder to construct a multi-quantile probabilistic forecasting model with only 85K parameters. Evaluated across 97 dataset configurations, the proposed model achieves state-of-the-art performance without fine-tuning, yielding a MASE of 0.808 while reducing the parameter count by 42% compared to TinyCast. This work effectively balances multi-domain generalization, predictive accuracy, and extreme compression efficiency.
📝 Abstract
Time series foundation models must preserve multi-domain breadth, probabilistic output, and multiple temporal scales, but parameter count grows when each scale receives a separate representation. We introduce Fracast-0, a probabilistic forecasting foundation model that exploits temporal self-similarity to reuse one operator across scales. A parameter-free detector extracts significant seasonal structure. The encoder applies a shared local block along a geometric dilation ladder with scale conditioning, while the decoder combines context-gathered states with an explicit seasonal future state and reuses a second block along another ladder before emitting nine quantiles. Pretraining across six corpora preserves multi-domain breadth within 85,001 parameters. On 97 GIFT-Eval configurations without per-dataset fine-tuning, Fracast-0 is the smallest of 28 evaluated checkpoints and remains non-dominated in the aggregate parameter-accuracy plane with MASE 0.808 and WQL 0.564. It uses 42.0% fewer parameters than TinyCast, whose MASE and WQL are 4.2% and 3.3% lower. These results support cross-scale weight reuse as a practical route to further time series foundation model compression.