🤖 AI Summary
This study addresses the limitation of existing attribution methods in multivariate forecasting, which fail to disentangle feature contributions to marginal uncertainty from those to dependency structures. To this end, it proposes a novel Entropy-Shapley game-theoretic hierarchical framework that integrates Shapley value theory, information entropy decomposition, and distribution regression. By leveraging conditional total correlation to explicitly characterize output dependencies, the framework isolates cross-component attribution terms to precisely identify the sources of uncertainty. This work bridges a critical theoretical gap in joint dependency attribution by successfully quantifying feature contributions to multivariate dependency structures. The effectiveness of the proposed approach is validated across diverse scenarios, including time-series foundation models.
📝 Abstract
Modern probabilistic machine learning models increasingly produce multivariate outputs with complex dependence structure, from multi-step time-series forecasts to sample path predictions. Understanding which input features drive the predictive uncertainty is important for risk-aware decisions, model diagnostics, and deciding whether the uncertainty should be mitigated or hedged against. This attribution problem requires a choice of how dependencies between output components are treated. Existing approaches reduce the output to a scalar through aggregation or projection before attribution, thereby obscuring whether features affect marginal uncertainty, dependence structure, or both, while component-wise analyses can miss dependence effects entirely. We close this gap by introducing a hierarchy of three entropy-based Shapley games that make this output-side choice explicit for any ordered multivariate outcome, ranging from per-component marginal entropy to fully joint entropy. The hierarchy isolates a cross-component attribution term that captures how each feature shifts the dependence between output components, a quantity invisible to component-wise methods. We establish a chain-rule decomposition of the joint attribution and characterize the cross-component term through conditional total correlation, providing both closed-form and sample-based estimators. Finally, we demonstrate how the framework captures differences in learned joint structure across probabilistic models from distributional regression to a zero-shot time series foundation model.