Scale-Invariant Training for Time Series Foundation Models
This study addresses the issue of "scale contamination" in time series foundation model training, where scale inversion causes high-magnitude samples to dominate gradients. To overcome this, we propose a scale-invariant training method employing Reversible Instance Normalization (ReVIN). For homogeneous loss functions such as mean squared error, the loss is computed directly on scaled targets, ensuring scale-independent optimization trajectories. This work provides the first theoretical proof and correction of biases in existing normalization schemes, establishing a unified loss paradigm that requires only a single line of code for integration. Experiments demonstrate that our approach consistently reduces the MASE metric across four architectures, yielding average improvements of approximately 20% on GIFT-Eval and M-competitions, thereby significantly enhancing forecasting accuracy and robustness.