Score
Designs, implements, and evaluates methods and pipelines that compute Shapley-based attributions (SHAP) for model predictions, producing local (instance-level) and aggregated (global) feature attributions and audit-ready explanation scores while satisfying properties such as local accuracy and consistency. Work includes efficient estimation and computation of Shapley values for different model and feature types, generating interpretable attribution outputs (e.g., per-feature or per-instance attributions), and analyzing and validating those attributions for model debugging, comparison, and accountability.
This work addresses the lack of a unified framework in existing feature attribution methods, which leads to opaque assumptions, incomparable results, and susceptibility to failure modes. The authors propose the first unified mathematical framework for locally additive attributions, systematically integrating Shapley values, path integrals, gradient-based methods, perturbation approaches, and CAM-style techniques through five core dimensions: value functions, reference points, paths, perturbation distributions, and conservation rules. Through axiomatic analysis, comparative matrices, and formal modeling, the study elucidates how attribution outcomes depend critically on underlying assumptions and establishes causal links between methodological choices and characteristic failure modes. To enhance rigor, the paper concludes with a ten-item reporting checklist designed to substantially improve the transparency, reproducibility, and reliability of attribution research.
Existing feature attribution methods (e.g., KernelSHAP, LIME) rely on global data distributions, leading to inaccurate characterization of local model behavior and distorted explanations. To address this, we propose VARSHAP—a model-agnostic local feature attribution method that introduces prediction variance reduction as the core Shapley value metric, the first such formulation. VARSHAP rigorously satisfies the efficiency, symmetry, and additivity axioms of Shapley values. It estimates conditional variances via Monte Carlo sampling, eliminating the need for surrogate models or distributional assumptions, and inherently exhibits robustness to data distribution shifts. Experiments on synthetic and real-world datasets demonstrate that VARSHAP improves attribution accuracy by 12–23% over KernelSHAP and LIME. Qualitative evaluations confirm its superior alignment with local decision logic, significantly mitigating the local explanation bias induced by global distribution dependence.
Existing feature attribution methods for learning-to-rank (LTR) lack ranking-aware theoretical foundations, often yielding contradictory or counterintuitive results that undermine interpretability. Method: This paper introduces the first game-theoretic, axiomatized framework for ranking—formally specifying ranking-specific axioms including ranking consistency and efficiency—and derives Rank-SHAP, the first axiomatic extension of Shapley values to ranking tasks. Contribution/Results: We evaluate Rank-SHAP on MSLR-WEB30K and Istella with state-of-the-art LTR models (e.g., LambdaMART, DeepRank) and validate it via user studies. Results demonstrate significant improvements in attribution consistency and alignment with human judgment. Axiomatic analysis further reveals that most existing attribution methods violate fundamental ranking axioms. This work establishes the first rigorous, axiom-based foundation for explainable LTR.
Existing Data Shapley methods require repeated retraining of data subsets, incurring prohibitive computational overhead, and yield generic contribution scores that cannot be tailored to specific target models. Method: We propose In-Run Data Shapley—the first framework enabling efficient, model-specific data contribution attribution for a single training run, including large language models. It embeds Shapley value theory directly into the dynamic parameter update process via gradient tracing and stochastic linear approximation, eliminating the need for auxiliary training. Contribution/Results: The method incurs negligible attribution overhead and supports fine-grained, pretraining-stage quantification of data value. Experiments demonstrate its interpretability and practical utility in copyright provenance and data curation. By bypassing iterative retraining, In-Run Data Shapley overcomes the computational bottleneck hindering data value assessment in large-scale models.
Existing feature attribution methods (e.g., SHAP) in information retrieval provide only document-level pointwise explanations, failing to capture inter-document relative ranking relationships within a ranked list. Method: This paper formally defines the list-level feature attribution problem and proposes a Shapley-value-based theoretical framework for joint attribution over entire ranking outputs. We introduce two novel evaluation paradigms to assess attribution correctness and completeness; identify contrastive decision-making as a fundamental constraint on attribution design; develop LTR-model-adapted attribution algorithms; and propose explanation-driven qualitative validation techniques. Results: Experiments on standard LTR benchmarks demonstrate that our method precisely identifies features governing relative document positioning, overcoming inherent limitations of selection-based explanations and significantly enhancing the interpretability of ranking models.
This study addresses the troubling inconsistency in feature attribution methods such as SHAP, which can yield substantially divergent explanations even for identical inputs and models, thereby undermining trustworthiness and auditability in high-stakes applications. The work formally defines and quantifies the phenomenon of “explanation multiplicity,” distinguishing its origins in model training or selection from inherent randomness in the explanation procedure itself. To assess stability, the authors introduce a dual-perspective metric incorporating both feature magnitude and ranking, establish a randomized null model as an interpretable baseline, and develop a comprehensive empirical evaluation framework spanning diverse datasets and model classes. Experiments demonstrate that explanation multiplicity is pervasive; relying solely on SHAP value magnitudes can lead to misleading conclusions, necessitating rank-sensitive metrics and principled baselines for reliable interpretability assessment.
This study addresses the strong model dependency of SHAP value interpretations, which lacks a standardized analytical framework and thereby limits reliable explanations of black-box model decisions in high-stakes applications. For the first time, this work systematically evaluates SHAP explanations across multiple mainstream machine learning models on diverse datasets, uncovering consistent patterns of model dependence. Furthermore, it proposes a generalized waterfall plot visualization method tailored for multi-class classification problems. Experimental results demonstrate the effectiveness and practical utility of the proposed approach, offering both theoretical grounding and actionable guidance for practitioners in the field of explainable artificial intelligence.
Traditional Shapley value methods struggle to simultaneously account for externalities among features and exogenous influences, leading to implausible explanations in complex causal structures. This work proposes DAG-SHAP, which introduces edge interventions into the Shapley attribution framework for the first time, treating edges—rather than nodes—as the fundamental units of attribution within a directed acyclic graph (DAG). This finer-grained approach enables more precise characterization of each feature’s role along causal pathways. To ensure scalability, we develop an efficient approximation algorithm and demonstrate through experiments on multiple real-world and synthetic datasets that DAG-SHAP achieves substantially improved attribution accuracy and interpretability compared to existing methods.
Traditional Shapley value computation is computationally prohibitive, and existing learnable explanation methods struggle with the non-uniform grids and irregular geometries commonly encountered in physical simulations. This work proposes OperatorSHAP—the first mesh-agnostic attribution method that extends Shapley values to function spaces. By integrating neural operator architectures with a learnable explainer, OperatorSHAP delivers consistent explanations across varying mesh resolutions without requiring model retraining. The method establishes a theoretical connection to the Aumann–Shapley value and demonstrates strong empirical alignment with discrete Shapley values across multiple grid resolutions. Consequently, it significantly enhances both the efficiency and generalization of model interpretability in physics-informed applications.
This work addresses the computational intractability of standard SHAP due to its #P-hard complexity in feature attribution by integrating causal knowledge into the interpretability framework. The authors propose Asymmetric Shapley Values (ASV) grounded in causal graphs, leveraging equivalence classes derived from topological orderings of the causal structure. They establish, for the first time, a polynomial-time exact algorithm for computing ASV under rooted directed tree structures and further develop an efficient approximation algorithm applicable to arbitrary causal directed acyclic graphs (DAGs). Experimental results demonstrate that the proposed approach substantially improves computational efficiency on real-world causal structures while preserving high-quality explanation fidelity.
This work addresses the theoretical shortcomings of prevailing post-hoc feature attribution methods—such as SHAP—which often lack formal guarantees and can mislead human decision-making in high-stakes settings. To overcome this limitation, the paper introduces a novel paradigm grounded in symbolic explainable artificial intelligence (XAI), leveraging formal, verifiable symbolic reasoning to reconstruct feature importance assignments. By integrating formal verification with feature attribution analysis, the proposed framework establishes a theoretically sound and certifiable approach to interpretability. This integration significantly enhances the reliability and trustworthiness of XAI in safety-critical applications, offering a rigorous and verifiable pathway for generating explanations in high-risk machine learning systems.