Score
Designs, implements, and analyzes methods that attribute a model’s output to individual inputs by performing input ablation (leave‑one‑out) experiments — removing or zeroing single inputs and measuring the change in model behavior. Produces per‑input attribution magnitudes and assesses attribution fidelity by comparing these LOO attributions to black‑box ablations or surrogate estimators.
This work addresses the lack of a unified framework in existing feature attribution methods, which leads to opaque assumptions, incomparable results, and susceptibility to failure modes. The authors propose the first unified mathematical framework for locally additive attributions, systematically integrating Shapley values, path integrals, gradient-based methods, perturbation approaches, and CAM-style techniques through five core dimensions: value functions, reference points, paths, perturbation distributions, and conservation rules. Through axiomatic analysis, comparative matrices, and formal modeling, the study elucidates how attribution outcomes depend critically on underlying assumptions and establishes causal links between methodological choices and characteristic failure modes. To enhance rigor, the paper concludes with a ten-item reporting checklist designed to substantially improve the transparency, reproducibility, and reliability of attribution research.
Existing data attribution methods (e.g., Leave-One-Out) perturb only individual training samples, ignoring collective interactions among samples and lacking a principled baseline mechanism—thus hindering counterfactual interpretation. This paper proposes Integrated Influence, the first method to incorporate both a controllable baseline dataset and a data degradation path into data attribution. It constructs a well-defined baseline, formalizes a sample-wise degradation sequence, and accumulates influence along this path using an integral-gradient-inspired approach—enabling global, counterfactual attribution for test predictions. Integrated Influence unifies mainstream influence-based methods (e.g., influence functions) under a theoretically grounded, computationally efficient framework. Experiments demonstrate its superiority in tasks such as mislabeled sample detection, achieving significantly higher stability, accuracy, and generalization compared to state-of-the-art baselines.
This study addresses the challenge of evaluating whether open-source language models can effectively serve as proxies to interpret the behavior of closed-source large language models when internal access is unavailable. Employing API-compatible methods—including log-odds probing, leave-one-out attribution, attention analysis, and input ablation—the authors conduct cross-model comparisons across 11 models spanning four major families: Llama, Qwen, GPT, and Gemini. Their findings reveal that predictive consistency substantially exceeds attribution consistency across model pairs. While white-box signals exhibit stability, they often fail to accurately reflect underlying causal mechanisms; in contrast, black-box input ablation more reliably captures the attribution behavior of closed-source models. These results uncover an “access-effectiveness inversion,” demonstrating that alignment in predictions alone is insufficient to support the transferability of mechanistic interpretations.
This paper addresses the challenge of attributing final model behavior to individual stages—such as pretraining, fine-tuning, and alignment—in multi-stage AI training. We propose the first *Accountable Attribution* framework to quantify the causal contribution of each stage to downstream model behavior. Methodologically, we develop an efficient, retraining-free estimator grounded in counterfactual reasoning and first-order optimization approximations, explicitly modeling training dynamics (e.g., learning rate, momentum, weight decay) and data distribution shifts across stages. Empirical evaluation across diverse tasks demonstrates that our framework accurately identifies the dominant training stage responsible for critical behavioral failures—including bias emergence and performance degradation. Our work provides an interpretable, computationally tractable tool for model debugging, trustworthy AI evaluation, and accountability assignment, thereby filling a key gap in causal analysis of AI training pipelines.
Existing climate attribution methods—such as optimal fingerprinting—exhibit limited performance under short temporal scales, low signal-to-noise ratios, or weak forcing–response relationships, and single-step modeling struggles to integrate heterogeneous climate information. This paper proposes a conditional multi-step attribution framework that, for the first time, formalizes the climate forcing–mediator–surface response pathway—as instantiated by stratospheric temperature and radiative flux—as an identifiable causal chain. The method integrates Bayesian inference with multivariate coupled response modeling to quantify forcing intensity. Robustness is enhanced via scalar response feature extraction and joint analysis of mediator variables. Applied to the 1991 Mount Pinatubo eruption, the framework substantially improves attribution confidence over conventional univariate temperature-based approaches, demonstrating its efficacy and novelty in high-noise, weak-signal regimes.
Existing attribution methods lack ground-truth causal explanations, rendering faithfulness evaluation unreliable. Method: This paper introduces BackX, a high-fidelity explainable AI benchmark that—uniquely—injects controllable causal attribution signals via backdoor triggers, rigorously satisfying faithfulness criteria including completeness, causality, and controllability. Contribution/Results: We provide theoretical guarantees that BackX outperforms both synthetic and real-world benchmarks in faithfulness assessment. We establish a standardized evaluation protocol incorporating attribution post-processing and cross-model consistency analysis. Empirically, BackX enables reproducible, high-discriminative evaluation across 12 state-of-the-art attribution methods, systematically exposing their causal faithfulness deficiencies. Furthermore, it inspires a novel attribution-based backdoor detection paradigm.
This study investigates whether the "abliteration" of refusal behavior in language models selectively removes refusals without altering other decision-making characteristics. Through a weekly stock price direction prediction task—conducted in a no-refusal setting—the authors systematically compare original and abliterated models using a frozen inference pipeline, decision tendency probes, and bootstrap confidence interval analysis to assess shifts in decision bias, confidence levels, and linguistic expression. The work reveals, for the first time, consistent and reproducible off-target decisional shifts across multiple model families: abliterated models uniformly exhibit greater optimism, produce more verbose explanations, and express less uncertainty, while changes in confidence vary by model architecture. These findings challenge the assumption that refusal can be precisely excised and further demonstrate that none of the examined models possess genuine economic forecasting ability.
Existing path-based feature attribution methods define trajectories in the input space, rendering them susceptible to path artifacts and unable to discern the semantic significance of input perturbations, which leads to unstable explanations. This work proposes Reveal-IG, a novel framework that lifts path attribution from the input space into a structured probe distribution space centered around the target sample, computing integrated gradients along distributional paths with respect to the model’s expected output. By supporting multi-scale image probes and modeling feature uncertainty in tabular data, Reveal-IG preserves attribution completeness while avoiding input-space artifacts. Experiments demonstrate that Reveal-IG produces stable, signed attributions on ImageNet classification and tabular regression tasks, significantly outperforming existing methods on sign-dependent metrics and remaining competitive on others, with synthetic diagnostics further confirming its robustness against artifacts.
This work addresses the failure of conventional single-component ablation-based attribution methods in Transformers, which overlook backup pathways due to the model’s self-repair mechanisms, often misclassifying critical components as irrelevant. To overcome this limitation, the authors propose Conditional Collaborative Ablation (CoAx), the first approach that formalizes self-repair as a conditional circuit completion task. CoAx quantifies the conditional increase in ablation effects across remaining units after removing a primary component, thereby uncovering latent second-order interactions and redundant pathways. The method operates in an unsupervised manner—requiring no labels and relying solely on model outputs—and integrates counterfactual patching, structured pruning, and cross-model transfer for precise attribution. Evaluated on the IOI task with GPT-2-small, CoAx improves AUC for backup head identification from 0.33 to 0.91, substantially outperforming baselines, and successfully generalizes across eight distinct models, enabling capability knockout and scalable pruning.