Semiparametric Inference for Conditional Shapley Feature Importance

📅 2026-09-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文针对特征依赖性问题,提出一种基于条件分布的半参数推断方法来估计Shapley特征重要性,并通过K折交叉拟合和U统计量校正平方损失以减少偏差。
📝 Abstract
Shapley values are widely used for post-hoc feature attribution, but most estimators return point quantities and do not quantify uncertainty, and popular implementations sample out-of-coalition features from their marginal distribution, which misattributes importance when features are dependent. This paper studies the conditional formulation, in which out-of-coalition features are integrated out under their true conditional distribution. The target is a global, loss-based importance that pairs a conditional value function with a SAGE-style loss aggregation. We propose a one-step estimator with K-fold cross-fitting and a U-statistic correction of the squared loss that removes the Monte Carlo bias of the naive plug-in; it is $\sqrt{n}$-consistent and asymptotically normal under double-robust rate conditions, and the resulting Wald interval attains nominal coverage. A Pinsker-type bound quantifies the bias from misspecifying the working copula class, while vine copulas keep conditional sampling tractable. In a Gaussian design study with n= 500, the empirical coverage of the 95% interval lies between 0.91 and 0.96 across all features, the test holds its Type-I rate at 0.05, and it reaches power one for moderate signals. Applied to the UCI Concrete and California Housing data, the method identifies the conditionally informative features with Bonferroni-controlled significance.
Problem

Research questions and friction points this paper is trying to address.

Semiparametric Inference
Conditional Shapley Feature Importance
Feature Attribution
Uncertainty Quantification
Dependent Features
Innovation

Methods, ideas, or system contributions that make the work stand out.

Conditional Shapley Feature Importance
K-fold Cross-fitting
U-statistic correction
Vine Copulas