From HL to H+L-1 Parameters: A Hankel-Toeplitz Forecaster for Long-Term Time Series Forecasting

šŸ“… 2026-09-27
šŸ“ˆ Citations: 0
✨ Influential: 0
šŸ“„ PDF
šŸ¤– AI Summary
This study addresses the parameter redundancy problem in linear models for long-term time series forecasting. Drawing upon classical stationary prediction theory, we propose the Hankel-Toeplitz Factorization (HTF) architecture. By exploiting the Hankel-Toeplitz structural properties of autocovariance matrices alongside innovation representations, HTF compresses full-rank prediction matrices into a compact form containing only H+Lāˆ’1 trainable coefficients, thereby enabling efficient parameter sharing. Evaluated across seven benchmark datasets, HTF achieves MSE forecasting accuracy comparable to dense linear models while reducing the parameter count by 75 to 229 times. This work establishes a new paradigm for long-sequence linear forecasting that is both theoretically grounded and highly parameter-efficient.
šŸ“ Abstract
Linear forecasters have shown competitive accuracy against Transformer-based models in long-term time series forecasting. We study how classical stationary prediction theory can guide parameter sharing for more compact linear forecasters. For centered second-order stationary processes with nonsingular history covariance, the minimum-MSE finite-window linear predictor factors into a Hankel cross-covariance matrix and an inverse Toeplitz covariance matrix. Shared lags and scale cancellation specify this predictor using $H+L-1$ autocorrelations for lookback $L$ and horizon $H$. Building on the innovations representation, our Hankel-Toeplitz Forecaster (HTF) learns one impulse response that defines both an inverse filter and a forecast map. We characterize the finite-history correction and, under summability assumptions, bound the excess risk of truncating the true filters. HTF uses $H+L-1$ trainable coefficients while allowing a full-rank forecasting matrix. Across seven benchmarks at $L=336$, its horizon-averaged MSE is within 1.2% of Dense Linear on each dataset with 75-229 times fewer trainable parameters.
Problem

Research questions and friction points this paper is trying to address.

long-term time series forecasting
linear forecasters
parameter efficiency
compact model
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hankel-Toeplitz Forecaster
Long-Term Time Series Forecasting
Parameter Sharing
Stationary Prediction Theory
Innovations Representation
šŸ”Ž Similar Papers
No similar papers found.