feature attribution analysis

Designs and implements algorithms, tools, and analyses that compute, compare, and visualize quantitative contributions of input features, internal units, or representations to model outputs — covering gradient- and propagation-based attributions, layer-wise relevance propagation, circuit- and feature-to-feature attribution, information-value and saliency measures, and extensions for embeddings and diffusion models. Builds attribution-based feature selection and importance-aware selection methods, constructs comparisons and ablations of attribution methods, and traces contribution flows through variables or pipeline steps to diagnose model behavior, dataset shortcuts, or component-level failures.

featureattributionanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$202K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Current AI interpretability methods are fragmented and lack a unified theoretical foundation. Method: This paper proposes the first unified analytical framework spanning three attribution paradigms—feature-, data-, and model-component-level attribution. By rigorously establishing the mathematical equivalence among perturbation analysis, gradient backpropagation, and linear approximations (e.g., Taylor expansions), it reveals their shared underlying mechanism: local sensitivity modeling. Contribution/Results: The framework standardizes terminology, aligns conceptual definitions, and unifies evaluation criteria—thereby significantly enhancing method interpretability, transferability, and reusability. It lowers entry barriers for newcomers while enabling advanced applications such as model editing, controllable steering, and AI governance. As a foundational contribution, it provides both theoretical grounding and practical scaffolding for next-generation interpretable AI systems.

AI Decision UnderstandingSimplificationUnified Method

AttributionScanner: A Visual Analytics System for Model Validation with Metadata-Free Slice Finding

Jan 12, 2024
XX
Xiwei Xuan
🏛️ University of California | Bosch | Splunk

Automatically identifying and attributing performance-deficient data slices (i.e., subpopulations) in unlabeled, unstructured image data remains challenging due to the absence of metadata and interpretable diagnostics. Method: This paper proposes a metadata-agnostic data slicing framework that jointly leverages gradient-based class attribution maps and clustering-driven slice discovery; introduces Attribution Mosaic—a novel visual analytics technique for slice-level attribution interpretation; and integrates a human-in-the-loop analysis pipeline with a plug-and-play attribution-consistency regularization mechanism for end-to-end model repair. Results: Evaluated on two benchmark vision datasets, the method achieves an average 5.2% improvement in slice-level accuracy, enables users to complete bias diagnosis and mitigation within 10 minutes, and reduces reliance on manual annotations by over 90%. Its core contributions are the first metadata-free, interpretable slice discovery method, attribution-driven visual analytics, and a deployable, end-to-end repair pipeline.

Detect and mitigate model issues like biases.Identify poor performance subgroups in datasets.Validate vision models without additional metadata.

Feature Attribution from First Principles

May 30, 2025
MT
Magamed Taimeskhanov
🏛️ Julius-Maximilians Universität Würzburg

Existing feature attribution methods lack rigorous theoretical foundations, and their empirical evaluation faces fundamental challenges. Method: We propose a first-principles–driven, bottom-up attribution framework: using indicator functions as atomic building blocks, attribution is recursively defined via function-space composition—bypassing ad hoc axiomatic constraints. Contribution/Results: We derive, for the first time, a closed-form attribution expression for deep ReLU networks, enabling efficient, differentiable computation. The framework unifies mainstream methods—including Integrated Gradients (IG) and DeepLIFT—under a common theoretical lens. Furthermore, it enables attribution-driven, differentiable evaluation objectives, facilitating end-to-end optimization of attribution quality. Combining theoretical rigor with practical scalability, this framework establishes a novel paradigm for interpretable AI.

Evaluating feature attribution methods empirically is challengingExisting axiomatic frameworks are often too restrictiveProposing a new feature attribution framework from first principles

Existing path-based feature attribution methods define trajectories in the input space, rendering them susceptible to path artifacts and unable to discern the semantic significance of input perturbations, which leads to unstable explanations. This work proposes Reveal-IG, a novel framework that lifts path attribution from the input space into a structured probe distribution space centered around the target sample, computing integrated gradients along distributional paths with respect to the model’s expected output. By supporting multi-scale image probes and modeling feature uncertainty in tabular data, Reveal-IG preserves attribution completeness while avoiding input-space artifacts. Experiments demonstrate that Reveal-IG produces stable, signed attributions on ImageNet classification and tabular regression tasks, significantly outperforming existing methods on sign-dependent metrics and remaining competitive on others, with synthetic diagnostics further confirming its robustness against artifacts.

attribution artifactsfeature attributioninput-space paths

Diffusion Attribution Score: Evaluating Training Data Influence in Diffusion Model

Oct 24, 2024
JL
Jinxu Lin
🏛️ The University of Sydney | City University of Hong Kong

To address copyright and privacy risks posed by training samples in diffusion models, this paper proposes the first direct data attribution method that quantifies the influence of individual training samples on the generated distribution. Unlike existing indirect paradigms relying on loss-function perturbations, we introduce the Distribution Attribution Score (DAS), a theoretically grounded metric defined via predictive distribution divergence—guaranteeing attribution consistency under mild assumptions. Our approach integrates statistical distance measures tailored to diffusion output distributions, gradient-based approximation for computational efficiency, and a scalable modeling framework enabling high-throughput attribution in large-scale models. Extensive experiments across multiple benchmarks and mainstream diffusion architectures—including Latent Diffusion Models (LDM) and Stable Diffusion (SD)—demonstrate substantial improvements over prior methods; notably, our LDM score establishes a new state-of-the-art. The implementation is publicly available.

Addressing misuse of copyrighted and private imagesEvaluating training data influence in diffusion modelsImproving accuracy of data attribution methods

Latest Papers

What's happening recently
View more

Nonparametric Data Attribution for Diffusion Models

Oct 15, 2025
YZ
Yutian Zhao
🏛️ Sea AI Lab | National University of Singapore | Singapore Management University

To address the limitations of existing data attribution methods for diffusion models—which typically require gradient computations or model retraining and thus hinder applicability in proprietary or large-scale settings—this paper proposes a gradient-free, retraining-free, nonparametric attribution method. Our approach leverages local patch-wise similarity between generated and training images, performing attribution via an analytically derived optimal scoring function in a multi-scale feature space. This constitutes the first natural extension of nonparametric attribution to multi-scale representations, without reliance on specific model architectures. By integrating convolutional acceleration and a purely data-driven framework, the method achieves both spatial interpretability and computational efficiency. Experiments demonstrate that our method attains attribution accuracy comparable to gradient-based approaches, significantly outperforms existing nonparametric baselines, and scales effectively to large datasets and real-world deployment scenarios.

Developing nonparametric attribution without gradients or retrainingMeasuring influence via patch-level image similarity analysisQuantifying training data influence on diffusion model outputs

This study addresses the limitation of existing data attribution methods for diffusion models, which overlook the dynamic evolution of semantics during generation. To this end, we propose CADT, a framework that constructs counterfactual trajectories to extend static scalar attribution into dynamic response analysis, thereby revealing the evolutionary mechanisms of training samples throughout the denoising process. Furthermore, this work pioneers a dynamic trajectory-based concept attribution paradigm that leverages covariance-aware kernel calibration to align query and training representations, enabling a deeper analytical shift from identifying "who influences" to understanding "how influence occurs." Extensive experiments demonstrate that CADT significantly outperforms existing baseline methods across hierarchical, compositional, and stylistic attribution tasks on multiple public datasets.

Concept AttributionCounterfactual TrajectoriesData Attribution

Automated Circuit Interpretation via Probe Prompting

Nov 10, 2025
GB
Giuseppe Birardi
🏛️ Orma Lab Srl

Neural network interpretability suffers from labor-intensive manual analysis of attribution maps—requiring ~2 hours per prompt—and lacks scalability. Method: We propose an automated subgraph extraction framework: (1) a concept probe generates high-impact features; (2) cross-prompt activation clustering and semantically aligned supernode construction identify recurrent structural motifs; and (3) three transparent, interpretability-driven rules—Semantic, Relationship, and Say-X—enable hierarchical semantic clustering. Results: Our method reveals a layered circuit structure: early Transformer layers host general-purpose mechanisms, while later layers specialize. Experiments show strong fidelity—0.83 average completeness and 0.762 activation similarity—while achieving a subgraph replacement score of 0.54 and concept-group consistency of 0.425, significantly outperforming geometric clustering baselines. The approach supports plug-and-play reproducibility, substantially improving both efficiency and trustworthiness of interpretability analysis.

Automating interpretation of neural network feature pathways to reduce manual analysis timeGrouping features by cross-prompt activation signatures to reveal behavioral coherenceTransforming attribution graphs into compact interpretable subgraphs using concept-aligned supernodes

Hot Scholars

WS

Wojciech Samek

Professor at TU Berlin, Head of AI Department at Fraunhofer HHI, BIFOLD Fellow
Deep LearningInterpretabilityExplainable AITrustworthy AI
RC

Ruoyu Chen

Institute of Information Engineering, Chinese Academy of Sciences.
Explainable AITrustworthy AIFoundation Model
XC

Xiaochun Cao

Sun Yat-sen University
Computer VisionArtificial IntelligenceMultimediaMachine Learning
DB

David Bau

Assistant Professor at Northeastern University
Machine LearningComputer VisionNLPSoftware Engineering
SL

Sebastian Lapuschkin

Head of Explainable AI, Fraunhofer Heinrich Hertz Institute
InterpretabilityExplainable AIXAIMachine Learning