Score
Designs and implements methods to compute and approximate asymmetric Shapley values (ASV), i.e., feature attributions that respect causal asymmetries and ordering in a causal graph; this includes building exact polynomial-time algorithms for special graph classes (for example rooted trees) and approximation algorithms for general directed acyclic graphs.
This work addresses the computational intractability of standard SHAP due to its #P-hard complexity in feature attribution by integrating causal knowledge into the interpretability framework. The authors propose Asymmetric Shapley Values (ASV) grounded in causal graphs, leveraging equivalence classes derived from topological orderings of the causal structure. They establish, for the first time, a polynomial-time exact algorithm for computing ASV under rooted directed tree structures and further develop an efficient approximation algorithm applicable to arbitrary causal directed acyclic graphs (DAGs). Experimental results demonstrate that the proposed approach substantially improves computational efficiency on real-world causal structures while preserving high-quality explanation fidelity.
Traditional Shapley value methods struggle to simultaneously account for externalities among features and exogenous influences, leading to implausible explanations in complex causal structures. This work proposes DAG-SHAP, which introduces edge interventions into the Shapley attribution framework for the first time, treating edges—rather than nodes—as the fundamental units of attribution within a directed acyclic graph (DAG). This finer-grained approach enables more precise characterization of each feature’s role along causal pathways. To ensure scalability, we develop an efficient approximation algorithm and demonstrate through experiments on multiple real-world and synthetic datasets that DAG-SHAP achieves substantially improved attribution accuracy and interpretability compared to existing methods.
Kernel methods suffer from poor interpretability, and exact Shapley value computation is typically intractable due to exponential time complexity. Method: This paper proposes PKeX-Shapley, the first algorithm enabling exact polynomial-time Shapley value computation under product kernel models. Its core innovation lies in exploiting the multiplicative structure of product kernels to derive a decomposable functional representation and a recursive Shapley value formula, integrating RKHS theory, functional space decomposition, and dynamic programming for efficiency. Contribution/Results: PKeX-Shapley reduces Shapley value computation complexity from exponential to polynomial time, achieving zero-approximation-error attribution in kernel regression and classification. Moreover, it generalizes to statistical discrepancy measures—including MMD and HSIC—enabling rigorous feature-level interpretability analysis while preserving theoretical fidelity to the underlying kernel model.
This work addresses the computational intractability of exact do-Shapley value estimation, which typically requires exponentially many interventions. The authors propose a novel reformulation that expresses the do-Shapley value as a function of irreducible sets in the causal graph, enabling nonparametric identification through single-element interventions alone. When the number of irreducible sets \( r \) is much smaller than \( 2^d \), this approach achieves exact computation in linear time. An adaptive estimator is developed that, with a query budget approaching \( r \), attains machine precision—outperforming existing methods by several orders of magnitude in accuracy. By leveraging this new representation based on irreducible sets, the method substantially reduces both computational complexity and identification burden, offering a scalable solution for causal attribution in high-dimensional settings.
Traditional Shapley value-based data valuation suffers from inaccuracy when applied to real-world datasets exhibiting heterogeneity and complex dependency structures, as it implicitly assumes data homogeneity and independence. Method: This paper proposes a structure-aware asymmetric Data Shapley framework. It introduces the asymmetry axiom—the first formal incorporation of asymmetry into data value quantification—thereby relaxing classical Shapley assumptions. Leveraging k-nearest neighbor graphs to capture intrinsic data structure, we design the first exact algorithm that simultaneously provides theoretical guarantees (e.g., fairness, efficiency, structure-awareness) and computational tractability. Contribution/Results: Extensive evaluation across diverse supervised learning tasks and data market scenarios demonstrates substantial improvements in contribution assessment accuracy and structural sensitivity. The open-source implementation has been widely adopted, establishing a new paradigm for data pricing and interpretable machine learning.
This work addresses the limitation of the classical Shapley value, which assumes contributor interchangeability and thus fails to capture dependencies or priority differences. We propose the Priority-Aware Shapley Value (PASV), the first framework that unifies hard precedence constraints and soft priority weights within a principled axiomatic foundation, generalizing existing approaches as special cases. To enable practical computation, we introduce “priority scanning” for sensitivity analysis and develop a Metropolis–Hastings sampler based on adjacent swaps, facilitating efficient Monte Carlo estimation under arbitrary priority structures and supporting asymptotic analysis under extreme weights. Experiments on data valuation (MNIST/CIFAR10) and feature attribution (Census Income) demonstrate that PASV yields fairer allocations aligned with real-world dependency structures, confirming its effectiveness and practical utility.
This work addresses the longstanding trade-off between accuracy and efficiency in feature- and node-level attribution for graph neural networks (GNNs), which typically rely on numerical approximations of path integrals. The authors propose APEX, a novel framework that introduces PolyGIN—a GNN architecture with a carefully designed polynomial form—enabling, for the first time, exact analytical computation of Aumann-Shapley attribution via path integrals. By integrating Gauss–Legendre quadrature with polynomial message passing, APEX guarantees that model outputs are bounded multivariate polynomials, thereby satisfying both completeness and computational efficiency in attribution. Empirical evaluations demonstrate that APEX maintains strong predictive performance across multiple graph benchmarks while achieving significantly higher attribution fidelity than existing baselines and drastically reducing the number of evaluation points required for path integration.
This work addresses the high computational complexity of Shapley value estimation in large-scale data valuation by modeling the utility function as a smooth functional of the mean embedding of empirical distributions in a reproducing kernel Hilbert space (RKHS). Leveraging tools from functional analysis and cooperative game theory, the authors conduct an asymptotic analysis that reveals how, as the number of data sources grows, the Shapley value is asymptotically characterized by a simple first-order dominant term. This finding elucidates the scaling behavior and structural properties of Shapley values in such settings. Building on this insight, the paper derives an interpretable and computationally tractable approximation, which not only provides a theoretical benchmark for existing estimation algorithms but also offers a principled foundation for efficient and reliable valuation in large-scale data markets.
This work addresses the limitations of Shapley values in identifying positive contributors under nonlinear evaluation metrics such as AUC, which stem from their inherent linearity assumption. To overcome this, the paper proposes a nonlinear attribution method that satisfies core axioms—including consistency, equal treatment, and efficiency—by leveraging an optimization framework inspired by the least core to approximate the utility function and yield a unique, optimal contribution vector. This approach transcends the conventional linear attribution paradigm, substantially enhancing attribution reliability while preserving essential axiomatic properties. Experimental results demonstrate that the proposed method outperforms Shapley variants that relax only the efficiency axiom, particularly on AUC-based evaluations, thereby validating the effectiveness and superiority of nonlinear attribution in cooperative settings.
本文解决了实时流数据中滑动窗口聚合值解释问题,通过为SUM、COUNT、AVG和方差提供闭式解的方法计算谓词级别的Shapley值。