Score
Designs and implements feature-ranking and subset-selection procedures that use SHAP (Shapley) attribution values produced by predictive models (including CatBoost/tree-SHAP variants) to score and rank input variables. Builds pipelines or algorithms that pick minimal predictive feature subsets and reduce dimensionality and redundancy so as to preserve or improve model generalization and predictive performance.
This work addresses the longstanding challenge of jointly optimizing model interpretability and predictive performance in machine learning. We propose a SHAP-driven regularized training framework, whose core innovation is the first direct incorporation of TreeSHAP attribution values into the loss function via a novel joint regularization term based on the entropy of the attribution distribution. This term simultaneously enforces sparsity, concentration, and cross-sample stability of feature importances. Unlike post-hoc methods, our approach is end-to-end trainable, applicable to mainstream tree-based models (e.g., XGBoost, LightGBM), and supports both regression and classification tasks. Extensive experiments across multiple benchmark datasets demonstrate that the proposed method improves model generalization, yields more robust and interpretable SHAP attributions, and maintains or exceeds baseline accuracy—without sacrificing predictive performance.
This study addresses the strong model dependency of SHAP value interpretations, which lacks a standardized analytical framework and thereby limits reliable explanations of black-box model decisions in high-stakes applications. For the first time, this work systematically evaluates SHAP explanations across multiple mainstream machine learning models on diverse datasets, uncovering consistent patterns of model dependence. Furthermore, it proposes a generalized waterfall plot visualization method tailored for multi-class classification problems. Experimental results demonstrate the effectiveness and practical utility of the proposed approach, offering both theoretical grounding and actionable guidance for practitioners in the field of explainable artificial intelligence.
This work addresses the lack of valid statistical inference for global SHAP value metrics, such as the $p$-th power mean. It establishes the first debiased estimation framework yielding asymptotically normal estimators: for $p \geq 2$, it employs U-statistics, while for $1 \leq p < 2$, it leverages Neyman orthogonal scores. To handle the non-smoothness of the target functional, a tunable temperature smoothing strategy is introduced. Furthermore, the SHAP curve is treated as a nuisance function, and its estimation is integrated within a semi-parametric inference framework coupled with empirical risk minimization. The resulting method provides reliable confidence intervals for global SHAP-based feature importance, thereby establishing a rigorous statistical foundation for model interpretation and feature selection.
Existing feature attribution methods for learning-to-rank (LTR) lack ranking-aware theoretical foundations, often yielding contradictory or counterintuitive results that undermine interpretability. Method: This paper introduces the first game-theoretic, axiomatized framework for ranking—formally specifying ranking-specific axioms including ranking consistency and efficiency—and derives Rank-SHAP, the first axiomatic extension of Shapley values to ranking tasks. Contribution/Results: We evaluate Rank-SHAP on MSLR-WEB30K and Istella with state-of-the-art LTR models (e.g., LambdaMART, DeepRank) and validate it via user studies. Results demonstrate significant improvements in attribution consistency and alignment with human judgment. Axiomatic analysis further reveals that most existing attribution methods violate fundamental ranking axioms. This work establishes the first rigorous, axiom-based foundation for explainable LTR.
Existing feature attribution methods (e.g., SHAP) in information retrieval provide only document-level pointwise explanations, failing to capture inter-document relative ranking relationships within a ranked list. Method: This paper formally defines the list-level feature attribution problem and proposes a Shapley-value-based theoretical framework for joint attribution over entire ranking outputs. We introduce two novel evaluation paradigms to assess attribution correctness and completeness; identify contrastive decision-making as a fundamental constraint on attribution design; develop LTR-model-adapted attribution algorithms; and propose explanation-driven qualitative validation techniques. Results: Experiments on standard LTR benchmarks demonstrate that our method precisely identifies features governing relative document positioning, overcoming inherent limitations of selection-based explanations and significantly enhancing the interpretability of ranking models.
This paper addresses the lack of statistical reliability in global feature importance assessment for black-box models. We propose φ-test, the first method that integrates Shapley-value-guided feature selection with selective inference to enable interpretable and statistically verifiable global feature selection and significance testing. φ-test constructs a linear surrogate model and outputs a global importance table containing Shapley values, regression coefficients, post-selection p-values, and confidence intervals. In regression tasks involving tree-based and neural network models, the selected features—typically few in number—retain over 90% of the original model’s predictive performance. Moreover, the selected feature sets exhibit high stability across different base models and bootstrap resamples. By bridging SHAP-based explanation with classical statistical inference, φ-test establishes a new paradigm for explainable AI that jointly ensures interpretability and statistical rigor.
This work addresses the challenge of feature selection in settings where features exhibit dependencies and their nonlinear relationships with the response are unknown, conditions under which conventional methods often fail to identify truly relevant features. The authors propose MinShap, a novel approach that replaces the average marginal contribution in Shapley values with the minimum marginal contribution to better isolate each feature’s direct effect. By integrating the faithfulness assumption of directed acyclic graphs (DAGs) with a multiple hypothesis testing framework, MinShap yields a theoretically grounded feature selection algorithm. Empirical evaluations on both synthetic and real-world datasets demonstrate that MinShap consistently outperforms established methods—including LOCO, GCM, and Lasso—particularly achieving higher accuracy and stability in small-sample regimes.
Existing SHAP algorithms suffer from low computational efficiency on tree ensemble models; mainstream approaches—Path-Dependent and Background SHAP—struggle to balance accuracy and scalability. This paper proposes WOODELF, the first unified SHAP framework based on pseudo-Boolean logic modeling, which jointly encodes decision tree structures and background data into efficiently solvable logical formulas. WOODELF is the first implementation to natively support Background SHAP, Path-Dependent SHAP, and Shapley/Banzhaf interaction values within a single, coherent framework. Implemented entirely in Python (NumPy/SciPy/CuPy), it requires no C++ or CUDA extensions. On a benchmark task with 3 million samples, 5 million background points, and 127 features, WOODELF achieves inference times of 162 seconds on CPU and 16 seconds on GPU—accelerating over state-of-the-art baselines by 16× to 165×. This advancement significantly enhances practicality and cross-platform compatibility for large-scale model interpretability.
This work addresses the lack of theoretical foundations for feature attribution in multi-output predictive models, particularly the unresolved question of whether Shapley values should be computed independently for each output. By extending classical cooperative game axioms—efficiency, symmetry, dummy player, and additivity—to the vector-valued setting, the paper establishes a rigidity theorem: any attribution rule satisfying these axioms must decompose as a component-wise sum across individual outputs. This result formally justifies the necessity of output-wise SHAP explanations. Empirical validation on biomedical benchmarks demonstrates that this component-wise approach not only preserves interpretability consistency but also substantially improves computational efficiency during both training and deployment of multi-output models.