interaction-aware attribution

Designs and implements methods that attribute model outputs to interactions among input variables by quantifying the contributions of feature subsets and multivariate combinations, including subset-level counterfactual utility estimates. Builds analyses and explanation artifacts that reveal coordinated patterns of features and their joint influence on predictions.

interaction-awareattribution

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.17
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Modeling and Discovering Direct Causes for Predictive Models

Dec 03, 2024
YC
Yizuo Chen
🏛️ University of California, Los Angeles | RTX Technology Research Center

This work addresses the problem of identifying input features that are direct causes of predictions in machine learning models—critical for optimizing data collection and evaluating model interpretability. To this end, we propose the first sound and complete algorithm for direct cause discovery, grounded in structural causal models (SCMs). Our method integrates conditional independence testing, constraint satisfaction solving, and a novel causal independence rule, substantially improving search efficiency. We provide formal theoretical guarantees establishing both correctness (soundness) and completeness. Empirical evaluation across diverse benchmark datasets demonstrates that our algorithm achieves significantly higher accuracy and computational efficiency in identifying direct causes compared to state-of-the-art alternatives. The results establish our approach as a robust, theoretically grounded tool for behavioral analysis of predictive models and causal-aware data engineering.

Developing algorithms for discovering direct causesIdentifying features directly causing predictionsModeling input-output behavior of predictive models

Explaining Models under Multivariate Bernoulli Distribution via Hoeffding Decomposition

Oct 08, 2025
BF
Baptiste Ferrere
🏛️ EDF R&D | Institut de Mathématiques de Toulouse

This work addresses the interpretability of predictive models—such as binary neural networks and Boolean networks—under multivariate Bernoulli inputs. Method: We propose an L² oblique projection analysis framework grounded in Hoeffding decomposition. Theoretically, we establish, for the first time, the explicit structure of Hoeffding decomposition under Bernoulli distributions: all higher-order interaction terms are orthogonal to the one-dimensional main subspace, enabling exact reverse engineering and closed-form solutions. This structure permits explicit derivation of global sensitivity metrics, including Sobol’ indices and Shapley effects. Computationally, the framework integrates Hoeffding decomposition, L² oblique projection, and variance attribution theory. Results: Numerical experiments demonstrate its effectiveness and scalability in high-dimensional, sparse binary input settings. To our knowledge, this is the first unified framework for model interpretation under discrete, finite-support inputs that simultaneously ensures theoretical rigor and computational feasibility.

Deriving explicit interpretability indicators for binary decision systemsExplaining predictive models with correlated Bernoulli inputs via decompositionExtending decomposition framework to finite countable input models

Counterfactual explainability of black-box prediction models

Nov 03, 2024
ZG
Zijun Gao
🏛️ University of Southern California | University of Cambridge

Existing interpretability methods predominantly rely on statistical associations, failing to uncover the causal mechanisms underlying black-box models—especially when input variables exhibit dependencies, leading to inaccurate attribution. This paper introduces “counterfactual interpretability,” a novel conceptual framework for causal attribution. It extends global sensitivity analysis to a counterfactual causal setting, establishing a complete algebraic system of explanations encompassing main effects, interaction effects, and variable dependency structures. By integrating functional ANOVA, Sobol indices, and DAG-guided causal sensitivity analysis, the framework delivers causal-driven, decomposable, and quantifiable interpretations for black-box models under arbitrary dependency structures. Experiments demonstrate that our method significantly outperforms mainstream association-based approaches on causal paradox benchmarks, validating its superiority in revealing true causal influences.

Develops counterfactual explainability for causal attribution in complex modelsEstimates causal mechanisms explaining income inequality by demographic factorsExtends global sensitivity analysis to dependent variables using causal graphs

A New Approach to Backtracking Counterfactual Explanations: A Causal Framework for Efficient Model Interpretability

May 05, 2025
PF
Pouria Fatemi
🏛️ Technical University of Munich | Ecole Polytechnique Federale de Lausanne (EPFL) | Sharif University of Technology

Traditional counterfactual explanations often neglect underlying causal structures, yielding unrealistic instances and incurring high computational overhead. To address this, we propose the Retrospective Causal Counterfactual (RCF) framework—the first to embed causal reasoning into a lightweight gradient-based backtracking search. RCF models the domain via a causal graph, enforces structural equation constraints, and adaptively selects intervention variables using local sensitivity analysis. Theoretically, RCF unifies multiple existing counterfactual methods under a single causal generalization bound; practically, it delivers strong actionability by generating interventions grounded in causal mechanisms. Empirically, on multiple benchmark datasets, RCF achieves an average 3.2× speedup over state-of-the-art methods while improving counterfactual realism by 41%. A user study further confirms its significantly enhanced operational utility.

Addresses computational cost in causal explanation methodsEnhances interpretability with causal counterfactual explanationsGeneralizes existing techniques for actionable model insights

Current AI interpretability methods are fragmented and lack a unified theoretical foundation. Method: This paper proposes the first unified analytical framework spanning three attribution paradigms—feature-, data-, and model-component-level attribution. By rigorously establishing the mathematical equivalence among perturbation analysis, gradient backpropagation, and linear approximations (e.g., Taylor expansions), it reveals their shared underlying mechanism: local sensitivity modeling. Contribution/Results: The framework standardizes terminology, aligns conceptual definitions, and unifies evaluation criteria—thereby significantly enhancing method interpretability, transferability, and reusability. It lowers entry barriers for newcomers while enabling advanced applications such as model editing, controllable steering, and AI governance. As a foundational contribution, it provides both theoretical grounding and practical scaffolding for next-generation interpretable AI systems.

AI Decision UnderstandingSimplificationUnified Method

Latest Papers

What's happening recently
View more

This study addresses the long-standing misconception that estimator ranking inconsistencies in data attribution stem from approximation errors, revealing instead that they originate from counterfactual norm mismatches. By formalizing influence as a counterfactual estimator, this work establishes norm analysis as a necessary prerequisite for comparing estimators. It derives local decompositions to analytically characterize signal interaction mechanisms, validated through linearized approximations and controlled experiments. The primary contribution is the first demonstration that behavioral proxy selection critically impacts attribution quality, proving that differing norms directly induce ranking discrepancies. Furthermore, the proposed behavior-aligned norm successfully identifies target samples overlooked by default methods, substantially improving attribution accuracy.

counterfactual specificationdata attributioninfluence estimation

This work addresses the limitations of existing explainable AI methods, which predominantly focus on associative predictions and fall short in supporting decision-making that requires causal reasoning and counterfactual analysis. To bridge this gap, the paper proposes a novel framework that integrates causal machine learning with intrinsically interpretable models—such as additive models and symbolic regression—by explicitly embedding causal inference mechanisms within the model architecture. This approach enables the explicit recovery of causal structures and functional forms among variables directly from cross-sectional data. While maintaining high predictive accuracy, the method achieves comprehensive transparency in system structure, causal relationships, and response mechanisms, thereby substantially enhancing both interpretability and causal reliability for trustworthy “What-if” analyses.

causal machine learningcausal relationshipsdecision support

Existing time series models lack effective methods for global interpretability, often limited to local explanations. This work introduces example-based global explanations to time series forecasting for the first time, proposing a model-agnostic, user-centered approach that automatically constructs representative subsets of time series by balancing sample importance and diversity through a domain-specific utility function, thereby generating concise and transparent summaries of global model behavior. Experiments demonstrate that the method efficiently produces high-quality summaries that are readily amenable to human evaluation. User studies further reveal that domain experts significantly prefer the model understanding provided by this approach, leading to a marked improvement in their comprehension of overall model behavior.

explainabilityglobal explanationsmodel interpretability

This study addresses the long-standing question of whether multivariate Kriging outperforms single-output modeling under heterotopic observations. The work establishes, for the first time, a theoretical link between output-specific design geometry and the estimability of cross-output dependencies, yielding a model-free diagnostic criterion. It derives an exact expression for predictive gain along with its geometric bounds, leading to a practical first-order net benefit criterion. Theoretical analysis reveals the statistical nonequivalence between collocated and heterotopic designs. Extensive experiments—conducted using separable multi-output Gaussian processes, radial basis functions, and linear models of coregionalization—validate the proposed criterion on synthetic datasets, an M/M/1 queueing system, and a multi-pollutant monitoring network, demonstrating both its effectiveness and practical utility.

design geometryGaussian processesheterotopic

Hot Scholars

QZ

Quanshi Zhang

Shanghai Jiao Tong University
Interpretable Machine Learning
JZ

Junpeng Zhang

Hebei Normal University
Information SecurityPrivacy-PreservingDifferential Privacy
QR

Qihan Ren

Shanghai Jiao Tong University
Explainable AIMachine LearningComputer VisionNatural Language Processing
CS

Chandan Singh

Senior researcher, Microsoft research
🔍 Interpretability🤖 Foundation models🧠 Neuroscience🌳 Transparent models