Score
Designs and implements algorithms, tools, and analyses that compute, compare, and visualize quantitative contributions of input features, internal units, or representations to model outputs — covering gradient- and propagation-based attributions, layer-wise relevance propagation, circuit- and feature-to-feature attribution, information-value and saliency measures, and extensions for embeddings and diffusion models. Builds attribution-based feature selection and importance-aware selection methods, constructs comparisons and ablations of attribution methods, and traces contribution flows through variables or pipeline steps to diagnose model behavior, dataset shortcuts, or component-level failures.
This work addresses the lack of a unified framework in existing feature attribution methods, which leads to opaque assumptions, incomparable results, and susceptibility to failure modes. The authors propose the first unified mathematical framework for locally additive attributions, systematically integrating Shapley values, path integrals, gradient-based methods, perturbation approaches, and CAM-style techniques through five core dimensions: value functions, reference points, paths, perturbation distributions, and conservation rules. Through axiomatic analysis, comparative matrices, and formal modeling, the study elucidates how attribution outcomes depend critically on underlying assumptions and establishes causal links between methodological choices and characteristic failure modes. To enhance rigor, the paper concludes with a ten-item reporting checklist designed to substantially improve the transparency, reproducibility, and reliability of attribution research.
Current AI interpretability methods are fragmented and lack a unified theoretical foundation. Method: This paper proposes the first unified analytical framework spanning three attribution paradigms—feature-, data-, and model-component-level attribution. By rigorously establishing the mathematical equivalence among perturbation analysis, gradient backpropagation, and linear approximations (e.g., Taylor expansions), it reveals their shared underlying mechanism: local sensitivity modeling. Contribution/Results: The framework standardizes terminology, aligns conceptual definitions, and unifies evaluation criteria—thereby significantly enhancing method interpretability, transferability, and reusability. It lowers entry barriers for newcomers while enabling advanced applications such as model editing, controllable steering, and AI governance. As a foundational contribution, it provides both theoretical grounding and practical scaffolding for next-generation interpretable AI systems.
Automatically identifying and attributing performance-deficient data slices (i.e., subpopulations) in unlabeled, unstructured image data remains challenging due to the absence of metadata and interpretable diagnostics. Method: This paper proposes a metadata-agnostic data slicing framework that jointly leverages gradient-based class attribution maps and clustering-driven slice discovery; introduces Attribution Mosaic—a novel visual analytics technique for slice-level attribution interpretation; and integrates a human-in-the-loop analysis pipeline with a plug-and-play attribution-consistency regularization mechanism for end-to-end model repair. Results: Evaluated on two benchmark vision datasets, the method achieves an average 5.2% improvement in slice-level accuracy, enables users to complete bias diagnosis and mitigation within 10 minutes, and reduces reliance on manual annotations by over 90%. Its core contributions are the first metadata-free, interpretable slice discovery method, attribution-driven visual analytics, and a deployable, end-to-end repair pipeline.
Existing feature attribution methods lack rigorous theoretical foundations, and their empirical evaluation faces fundamental challenges. Method: We propose a first-principles–driven, bottom-up attribution framework: using indicator functions as atomic building blocks, attribution is recursively defined via function-space composition—bypassing ad hoc axiomatic constraints. Contribution/Results: We derive, for the first time, a closed-form attribution expression for deep ReLU networks, enabling efficient, differentiable computation. The framework unifies mainstream methods—including Integrated Gradients (IG) and DeepLIFT—under a common theoretical lens. Furthermore, it enables attribution-driven, differentiable evaluation objectives, facilitating end-to-end optimization of attribution quality. Combining theoretical rigor with practical scalability, this framework establishes a novel paradigm for interpretable AI.
Existing path-based feature attribution methods define trajectories in the input space, rendering them susceptible to path artifacts and unable to discern the semantic significance of input perturbations, which leads to unstable explanations. This work proposes Reveal-IG, a novel framework that lifts path attribution from the input space into a structured probe distribution space centered around the target sample, computing integrated gradients along distributional paths with respect to the model’s expected output. By supporting multi-scale image probes and modeling feature uncertainty in tabular data, Reveal-IG preserves attribution completeness while avoiding input-space artifacts. Experiments demonstrate that Reveal-IG produces stable, signed attributions on ImageNet classification and tabular regression tasks, significantly outperforming existing methods on sign-dependent metrics and remaining competitive on others, with synthetic diagnostics further confirming its robustness against artifacts.
To address copyright and privacy risks posed by training samples in diffusion models, this paper proposes the first direct data attribution method that quantifies the influence of individual training samples on the generated distribution. Unlike existing indirect paradigms relying on loss-function perturbations, we introduce the Distribution Attribution Score (DAS), a theoretically grounded metric defined via predictive distribution divergence—guaranteeing attribution consistency under mild assumptions. Our approach integrates statistical distance measures tailored to diffusion output distributions, gradient-based approximation for computational efficiency, and a scalable modeling framework enabling high-throughput attribution in large-scale models. Extensive experiments across multiple benchmarks and mainstream diffusion architectures—including Latent Diffusion Models (LDM) and Stable Diffusion (SD)—demonstrate substantial improvements over prior methods; notably, our LDM score establishes a new state-of-the-art. The implementation is publicly available.
To address the limitations of existing data attribution methods for diffusion models—which typically require gradient computations or model retraining and thus hinder applicability in proprietary or large-scale settings—this paper proposes a gradient-free, retraining-free, nonparametric attribution method. Our approach leverages local patch-wise similarity between generated and training images, performing attribution via an analytically derived optimal scoring function in a multi-scale feature space. This constitutes the first natural extension of nonparametric attribution to multi-scale representations, without reliance on specific model architectures. By integrating convolutional acceleration and a purely data-driven framework, the method achieves both spatial interpretability and computational efficiency. Experiments demonstrate that our method attains attribution accuracy comparable to gradient-based approaches, significantly outperforms existing nonparametric baselines, and scales effectively to large datasets and real-world deployment scenarios.
This study addresses the limitation of existing data attribution methods for diffusion models, which overlook the dynamic evolution of semantics during generation. To this end, we propose CADT, a framework that constructs counterfactual trajectories to extend static scalar attribution into dynamic response analysis, thereby revealing the evolutionary mechanisms of training samples throughout the denoising process. Furthermore, this work pioneers a dynamic trajectory-based concept attribution paradigm that leverages covariance-aware kernel calibration to align query and training representations, enabling a deeper analytical shift from identifying "who influences" to understanding "how influence occurs." Extensive experiments demonstrate that CADT significantly outperforms existing baseline methods across hierarchical, compositional, and stylistic attribution tasks on multiple public datasets.
本文通过引入概念定向归因(CTA)方法,解决了线性探针如何产生内部概念表示的问题,并解释了影响探针性能的内部计算机制。
Neural network interpretability suffers from labor-intensive manual analysis of attribution maps—requiring ~2 hours per prompt—and lacks scalability. Method: We propose an automated subgraph extraction framework: (1) a concept probe generates high-impact features; (2) cross-prompt activation clustering and semantically aligned supernode construction identify recurrent structural motifs; and (3) three transparent, interpretability-driven rules—Semantic, Relationship, and Say-X—enable hierarchical semantic clustering. Results: Our method reveals a layered circuit structure: early Transformer layers host general-purpose mechanisms, while later layers specialize. Experiments show strong fidelity—0.83 average completeness and 0.762 activation similarity—while achieving a subgraph replacement score of 0.54 and concept-group consistency of 0.425, significantly outperforming geometric clustering baselines. The approach supports plug-and-play reproducibility, substantially improving both efficiency and trustworthiness of interpretability analysis.
LumiXAI解决了解释模型时软件碎片化问题,通过提供一个模块化的全栈框架,整合了特征归因分析,并支持多种用户访问。