Time Series Forecasting Benchmarks Need Scenario-Grounded Stress Testing

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of existing time series forecasting benchmarks, which evaluate a narrow scope and fail to capture structured anomalies or system-level failures in real-world scenarios. This challenge is further compounded by the unauditable pretraining data of foundation models, which impedes rigorous generalization assessment. To overcome these issues, this work proposes a scenario-based stress testing framework that transcends conventional noise perturbation by incorporating causal and system-level failure modeling. By integrating semantic scenarios, explicit failure operators, and graded difficulty levels, the framework jointly evaluates historical inputs, future targets, and multidimensional failure metrics. Ultimately, this research establishes a novel paradigm combining attribution capability with deployment relevance, effectively revealing model failure mechanisms under specific operational conditions.
📝 Abstract
Time series forecasting (TSF) increasingly drives decisions in transportation, energy, finance, healthcare, and infrastructure, yet current evaluation remains overly narrow: standard benchmarks reward low held-out error, while robustness studies typically reduce failure to Gaussian noise, random masking, or bounded adversarial perturbations. This obscures the real failure modes of deployed forecasting systems. Input-side anomalies are not merely noisier inputs: they often reflect structured events that alter temporal dynamics, break cross-variable dependencies, induce regime shifts, or propagate from faulty sensors to downstream decisions. These semantic, causal, and system-level failures cannot be faithfully captured by i.i.d. perturbations alone. The rise of TSF foundation models makes this evaluation gap more urgent, as unauditable pretraining corpora make held-out generalization increasingly unreliable. We therefore advocate scenario-grounded stress testing. Each test instance should include historical inputs and future targets, together with a semantic scenario, an explicit failure operator, and a measurable difficulty level. This shift makes evaluation interpretable, attributable, and deployment-relevant and friendly, enabling the community to ask not only which model is accurate, but under what conditions it fails and why.
Problem

Research questions and friction points this paper is trying to address.

Time Series Forecasting
Benchmark Evaluation
Stress Testing
Robustness
Foundation Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Time Series Forecasting
Scenario-Grounded Stress Testing
Failure Operators
Robustness Evaluation
Foundation Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.