š¤ AI Summary
This study addresses the parameter redundancy problem in linear models for long-term time series forecasting. Drawing upon classical stationary prediction theory, we propose the Hankel-Toeplitz Factorization (HTF) architecture. By exploiting the Hankel-Toeplitz structural properties of autocovariance matrices alongside innovation representations, HTF compresses full-rank prediction matrices into a compact form containing only H+Lā1 trainable coefficients, thereby enabling efficient parameter sharing. Evaluated across seven benchmark datasets, HTF achieves MSE forecasting accuracy comparable to dense linear models while reducing the parameter count by 75 to 229 times. This work establishes a new paradigm for long-sequence linear forecasting that is both theoretically grounded and highly parameter-efficient.
š Abstract
Linear forecasters have shown competitive accuracy against Transformer-based models in long-term time series forecasting. We study how classical stationary prediction theory can guide parameter sharing for more compact linear forecasters. For centered second-order stationary processes with nonsingular history covariance, the minimum-MSE finite-window linear predictor factors into a Hankel cross-covariance matrix and an inverse Toeplitz covariance matrix. Shared lags and scale cancellation specify this predictor using $H+L-1$ autocorrelations for lookback $L$ and horizon $H$. Building on the innovations representation, our Hankel-Toeplitz Forecaster (HTF) learns one impulse response that defines both an inverse filter and a forecast map. We characterize the finite-history correction and, under summability assumptions, bound the excess risk of truncating the true filters. HTF uses $H+L-1$ trainable coefficients while allowing a full-rank forecasting matrix. Across seven benchmarks at $L=336$, its horizon-averaged MSE is within 1.2% of Dense Linear on each dataset with 75-229 times fewer trainable parameters.