Beyond Model Ranking: Regime Diagnosis for Distributional-Statistical Misspecification in Industrial Time-Series Forecasting

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the systematic bias and train-deployment gap in industrial time series forecasting caused by mismatches between loss function priors and data distributions. We propose a diagnostic framework based on regime-wise relative bias vectors (RBV) that decomposes prediction bias into a model-agnostic intrinsic lower bound and a training-attributable excess component. This approach reveals distributional shape biases obscured by aggregate metrics and establishes an affine relationship between risk and pathology-mixed weights to distinguish optimization-driven from bias-driven failures. Experiments on RetailShiftBench and M5 benchmarks demonstrate that regime-aware training effectively eliminates pooling-induced bias, yielding gains superior to mere capacity scaling. These findings provide mechanism-oriented theoretical guidance for model selection and loss design.
📝 Abstract
Time-series forecasting models achieve strong benchmark performance but exhibit severe systematic bias in industrial deployments. This train--deploy gap is conventionally attributed to temporal-structural errors or distribution shifts. We characterize a complementary source that these explanations overlook: canonical losses embed fixed statistical priors, while industrial demand mixes benign and pathological regimes---zero-inflation, skewness, high variability---in which these priors are systematically violated. The induced bias persists even under perfect temporal modeling, remains in a distributional-shape component that normalization cannot remove, and creates an aggregation trade-off invisible to aggregate metrics. We turn these observations into an evaluation toolkit centered on the Regime-wise Relative Bias Vector (RBV): a metric-agnostic, regime-decomposed diagnostic that audits how pooled training allocates systematic mismatch across pathological subpopulations. A controlled attribution analysis decomposes RBV into a model-independent intrinsic floor, set by each loss's estimand, and an excess component attributable to training, tracing observed bias to the loss rather than the model. A large-scale study---13 loss objectives, 3 seeds, 60,000+ series spanning RetailShiftBench and M5, with random-split controls---shows that regime-aware diagnosis separates optimization-type from bias-type failure, and that regime-aware training resolves the pooling-induced bias that capacity scaling cannot, for mean-type losses. A formal structural observation, that risk under evaluation-distribution contamination is affine in the pathology mixture weight, grounds these findings. Our work complements model ranking with mechanism-grounded, regime-oriented evaluation.
Problem

Research questions and friction points this paper is trying to address.

time-series forecasting
distributional-statistical misspecification
systematic bias
pathological regimes
loss function
Innovation

Methods, ideas, or system contributions that make the work stand out.

Regime-wise Relative Bias Vector
Distributional-Statistical Misspecification
Loss Function Priors
Controlled Attribution Analysis
Time-Series Forecasting
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Pengyu Nie
Pengyu Nie
University of Waterloo
Software EngineeringNatural Language ProcessingProgramming Languages
C
Chenglang Xu
JD.com, Inc.
Y
Yaoshi Chen
JD.com, Inc.
C
Chaogan Ren
JD.com, Inc.
W
Wei Hu
JD.com, Inc.
C
Chao Yang
JD.com, Inc.
J
Jiangong Zhang
JD.com, Inc.