A Unified Scaling Law for Time Series Foundation Models

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear interplay between model capacity and historical information in time series foundation models, as well as its impact on predictive capability. Through large-scale empirical analysis, Gaussian regression theory, and activation intervention experiments, this work establishes a unified scaling law that quantifies the nonlinear relationships among capacity, history length, and prediction horizon, further distilled into a concise five-parameter formulation. The research reveals distinct information utilization mechanisms between full fine-tuning and frozen models. Notably, it accurately predicts the performance of Toto 2.0 with an error margin below 1.5% and formalizes a paradigm for reusing historical scaling rules. Collectively, these contributions provide a rigorous theoretical foundation for the efficient scaling of time series foundation models.
📝 Abstract
We develop a Unified Scaling Law and a Unified Theory of Time Series Learning to understand how model capacity and historical information support forecasting. Across different lookback lengths and forecast horizons, we analyze 18,768 experimental cells from 21 checkpoints on 23 dataset-frequency tasks spanning six domains. Our empirical methodology integrates local resource relations into a parsimonious, fitted five-parameter law: capacity gains increase with history, context gains diminish toward saturation, and horizon effects enter as a common shift. Fitted without Toto 2.0, the law predicts its horizon-averaged capacity-scaling curves with mean absolute percentage errors of 1.09% and 1.50% at input lengths 2048 and 4096. To understand how history supports prediction, our learning theory uses Gaussian regression to analyze rule identification and predictive capability. We hypothesize that full-shot models learn by accumulating information in weights, while frozen time series foundation models (TSFMs) use history by extracting information through activations. Matched-history comparisons establish the predictive value of additional history. Controlled parameter exchanges and activation interventions provide evidence that history-derived rule information can be retained, reused across queries, and used to recover a contribution to long-context prediction. Together, these findings inform capacity scaling, context allocation, and the development of models that retain and apply historical rules. Code and main results are available at https://github.com/Fifthky/UniScale.
Problem

Research questions and friction points this paper is trying to address.

Time Series Foundation Models
Scaling Law
Forecasting
Model Capacity
Context Length
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified Scaling Law
Time Series Foundation Models
Gaussian Regression
Activation Intervention
Context Allocation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.