🤖 AI Summary
This study addresses the limitations of existing time series forecasting benchmarks, which evaluate a narrow scope and fail to capture structured anomalies or system-level failures in real-world scenarios. This challenge is further compounded by the unauditable pretraining data of foundation models, which impedes rigorous generalization assessment. To overcome these issues, this work proposes a scenario-based stress testing framework that transcends conventional noise perturbation by incorporating causal and system-level failure modeling. By integrating semantic scenarios, explicit failure operators, and graded difficulty levels, the framework jointly evaluates historical inputs, future targets, and multidimensional failure metrics. Ultimately, this research establishes a novel paradigm combining attribution capability with deployment relevance, effectively revealing model failure mechanisms under specific operational conditions.
📝 Abstract
Time series forecasting (TSF) increasingly drives decisions in transportation, energy, finance, healthcare, and infrastructure, yet current evaluation remains overly narrow: standard benchmarks reward low held-out error, while robustness studies typically reduce failure to Gaussian noise, random masking, or bounded adversarial perturbations. This obscures the real failure modes of deployed forecasting systems. Input-side anomalies are not merely noisier inputs: they often reflect structured events that alter temporal dynamics, break cross-variable dependencies, induce regime shifts, or propagate from faulty sensors to downstream decisions. These semantic, causal, and system-level failures cannot be faithfully captured by i.i.d. perturbations alone. The rise of TSF foundation models makes this evaluation gap more urgent, as unauditable pretraining corpora make held-out generalization increasingly unreliable. We therefore advocate scenario-grounded stress testing. Each test instance should include historical inputs and future targets, together with a semantic scenario, an explicit failure operator, and a measurable difficulty level. This shift makes evaluation interpretable, attributable, and deployment-relevant and friendly, enabling the community to ask not only which model is accurate, but under what conditions it fails and why.