🤖 AI Summary
This study addresses the phenomenon wherein time series models with comparable performance produce divergent prediction trajectories, a problem overlooked by existing research that focuses solely on pointwise errors. To this end, this work proposes the "temporal predictive multiplicity" framework, which is the first to formally define predictive multiplicity from a complete trajectory perspective. Theoretically, it demonstrates that single-timepoint constraints are insufficient to eliminate trajectory-level divergence, thereby bridging a critical gap in the literature. Extensive experiments encompassing 19 neural architectures and 11 datasets confirm that near-optimal models exhibit significant trajectory variation, and that trajectory-level and pointwise divergences are mutually independent. These findings reveal potential risks of such discrepancies for downstream applications.
📝 Abstract
Models with near-identical predictive performance can yield substantially different predictions, a phenomenon known as predictive multiplicity. Prior work has mostly studied this at the level of individual scalar outputs. In time-series forecasting, however, predictions across horizons jointly define a trajectory, and horizon-wise comparisons can hide important differences in predictive behavior. To address this problem, we introduce temporal predictive multiplicity, a framework that characterizes disagreement over complete forecast trajectories among models with near-identical predictive performance. We show that constraining predictive performance alone can still admit a broad range of different trajectories. We further show that constraining multiplicity at individual horizons partially reduces, but does not eliminate, trajectory-level multiplicity. Experiments with 19 neural forecasting architectures on 11 datasets confirm that near-optimal models can exhibit substantial variability in the forecast trajectories they produce, and trajectory-level disagreement is largely unrelated to horizon-wise disagreement. Our framework, therefore, exposes a gap in existing multiplicity studies: models with indistinguishable predictive performance imply fundamentally different temporal trajectories, with consequential downstream effects.