🤖 AI Summary
This study addresses the challenge that covariate effects vary dynamically across contexts, hindering traditional forecasting methods from accurately assessing their influence and extracting effective feedback from errors. To this end, we propose an experience-based forecasting framework in which a frozen time-series foundation model generates baseline predictions, while a large language model performs judgmental adjustments by integrating contextual information with historical experience. The core innovations include an explicit covariate judgment mechanism that decouples effect assessment from numerical adjustment, and a residual-guided experience reconstruction strategy that substitutes raw judgments with observed residuals to accumulate reliable experience. Extensive experiments on multiple real-world datasets demonstrate that the proposed approach outperforms strong baselines, validating both the effectiveness of explicit judgment in enhancing adjustment accuracy and the reliability of residual-guided experience accumulation.
📝 Abstract
Covariate effects vary across contexts and shift over time, requiring forecasters to assess how to use them for each forecasting context. As forecasting proceeds, observations for earlier forecasts become available, providing feedback on past covariate use for subsequent forecasts. However, when multiple covariates act together, the forecast error reveals the numerical discrepancy from the observation but not how the covariates should have been used. We introduce JudgeCast, an experience-based framework for time series forecasting with covariates. Following the judgmental adjustment practice, a frozen TSFM provides the base forecast, while a frozen LLM uses the current context and relevant experience to adjust it. Within the adjustment, assessing covariate effects and determining the numerical adjustment serve distinct roles, so JudgeCast first forms explicit covariate-wise judgments and then determines the adjustment. After observation, JudgeCast uses the observed residual of the base forecast to reconstruct alternative judgments and evaluates the original and alternatives through their resulting adjustments. The best-performing decision is selected and retained as validated experience for subsequent forecasts. Across diverse real-world datasets, JudgeCast outperforms strong baselines. Ablations show that explicit covariate-wise judgment can improve forecast-time adjustment, while residual-guided experience construction yields more reliable forecasting gains than retaining raw decisions as experience.