๐ค AI Summary
This work addresses the common oversight in traditional time series forecasting, where evaluation focuses predominantly on point accuracy while neglecting temporal consistencyโi.e., the stability of predictions for the same future timestamp when issued from different origins. To remedy this, the authors propose the Accuracy and Consistency Score (AC Score), a differentiable, user-weighted evaluation metric that explicitly incorporates stability into the assessment of multi-step probabilistic forecasts and serves as an end-to-end training objective. By optimizing the AC Score within a seasonal ARIMA framework, experiments on the M4 Hourly dataset demonstrate that the approach achieves comparable or superior point forecast accuracy while reducing prediction volatility for identical target timestamps by 75%.
๐ Abstract
Traditional time series forecasting methods optimize for accuracy alone. This objective neglects temporal consistency, in other words, how consistently a model predicts the same future event as the forecast origin changes. We introduce the forecast accuracy and coherence score (forecast AC score for short) for measuring the quality of probabilistic multi-horizon forecasts in a way that accounts for both multi-horizon accuracy and stability. Our score additionally allows user-specified weights to balance accuracy and consistency requirements. As an example application, we implement the score as a differentiable objective function for training seasonal auto-regressive integrated models and evaluate it on the M4 Hourly benchmark dataset. Results demonstrate substantial improvements over traditional maximum likelihood estimation. Regarding stability, the AC-optimized model generated out-of-sample forecasts with 91.1\% reduced vertical variance relative to the MLE-fitted model. In terms of accuracy, the AC-optimized model achieved considerable improvements for medium-to-long-horizon forecasts. While one-step-ahead forecasts exhibited a 7.5\% increase in MAPE, all subsequent horizons experienced an improved accuracy as measured by MAPE of up to 26\%. These results indicate that our metric successfully trains models to produce more stable and accurate multi-step forecasts in exchange for some degradation in one-step-ahead performance.