compute leave-one-out attributions

Designs, implements, and analyzes methods that attribute a model’s output to individual inputs by performing input ablation (leave‑one‑out) experiments — removing or zeroing single inputs and measuring the change in model behavior. Produces per‑input attribution magnitudes and assesses attribution fidelity by comparing these LOO attributions to black‑box ablations or surrogate estimators.

computeleave-one-outattributions

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.34
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Integrated Influence: Data Attribution with Baseline

Aug 07, 2025
LY
Linxiao Yang
🏛️ DAMO Academy | Alibaba Group

Existing data attribution methods (e.g., Leave-One-Out) perturb only individual training samples, ignoring collective interactions among samples and lacking a principled baseline mechanism—thus hindering counterfactual interpretation. This paper proposes Integrated Influence, the first method to incorporate both a controllable baseline dataset and a data degradation path into data attribution. It constructs a well-defined baseline, formalizes a sample-wise degradation sequence, and accumulates influence along this path using an integral-gradient-inspired approach—enabling global, counterfactual attribution for test predictions. Integrated Influence unifies mainstream influence-based methods (e.g., influence functions) under a theoretically grounded, computationally efficient framework. Experiments demonstrate its superiority in tasks such as mislabeled sample detection, achieving significantly higher stability, accuracy, and generalization compared to state-of-the-art baselines.

Addressing lack of baseline flexibility in attribution explanationsImproving reliability of data attribution and mislabel identificationOvercoming local-based limitations in data attribution methods

This study addresses the challenge of evaluating whether open-source language models can effectively serve as proxies to interpret the behavior of closed-source large language models when internal access is unavailable. Employing API-compatible methods—including log-odds probing, leave-one-out attribution, attention analysis, and input ablation—the authors conduct cross-model comparisons across 11 models spanning four major families: Llama, Qwen, GPT, and Gemini. Their findings reveal that predictive consistency substantially exceeds attribution consistency across model pairs. While white-box signals exhibit stability, they often fail to accurately reflect underlying causal mechanisms; in contrast, black-box input ablation more reliably captures the attribution behavior of closed-source models. These results uncover an “access-effectiveness inversion,” demonstrating that alignment in predictions alone is insufficient to support the transferability of mechanistic interpretations.

attributionclosed language modelsmechanistic interpretability

Accountability Attribution: Tracing Model Behavior to Training Processes

May 30, 2025
SZ
Shichang Zhang
🏛️ Harvard University | University of California, Los Angeles | University of Illinois Urbana-Champaign

This paper addresses the challenge of attributing final model behavior to individual stages—such as pretraining, fine-tuning, and alignment—in multi-stage AI training. We propose the first *Accountable Attribution* framework to quantify the causal contribution of each stage to downstream model behavior. Methodologically, we develop an efficient, retraining-free estimator grounded in counterfactual reasoning and first-order optimization approximations, explicitly modeling training dynamics (e.g., learning rate, momentum, weight decay) and data distribution shifts across stages. Empirical evaluation across diverse tasks demonstrates that our framework accurately identifies the dominant training stage responsible for critical behavioral failures—including bias emergence and performance degradation. Our work provides an interpretable, computationally tractable tool for model debugging, trustworthy AI evaluation, and accountability assignment, thereby filling a key gap in causal analysis of AI training pipelines.

Identify accountable stages for model behaviorsQuantify stage effects without retraining modelsTrace model behavior to specific training stages

Conditional multi-step attribution for climate forcings

Sep 02, 2024
CW
C. Wentland
🏛️ Sandia National Laboratories | Zicklin School of Business | Baruch College | CUNY

Existing climate attribution methods—such as optimal fingerprinting—exhibit limited performance under short temporal scales, low signal-to-noise ratios, or weak forcing–response relationships, and single-step modeling struggles to integrate heterogeneous climate information. This paper proposes a conditional multi-step attribution framework that, for the first time, formalizes the climate forcing–mediator–surface response pathway—as instantiated by stratospheric temperature and radiative flux—as an identifiable causal chain. The method integrates Bayesian inference with multivariate coupled response modeling to quantify forcing intensity. Robustness is enhanced via scalar response feature extraction and joint analysis of mediator variables. Applied to the 1991 Mount Pinatubo eruption, the framework substantially improves attribution confidence over conventional univariate temperature-based approaches, demonstrating its efficacy and novelty in high-noise, weak-signal regimes.

Attributing climate impacts to natural and anthropogenic forcingsEnhancing certainty using physical pathways and multivariate dataImproving attribution in low signal-to-noise regimes

Backdoor-based Explainable AI Benchmark for High Fidelity Evaluation of Attribution Methods

May 02, 2024
PY
Peiyu Yang
🏛️ The University of Western Australia | The University of Melbourne

Existing attribution methods lack ground-truth causal explanations, rendering faithfulness evaluation unreliable. Method: This paper introduces BackX, a high-fidelity explainable AI benchmark that—uniquely—injects controllable causal attribution signals via backdoor triggers, rigorously satisfying faithfulness criteria including completeness, causality, and controllability. Contribution/Results: We provide theoretical guarantees that BackX outperforms both synthetic and real-world benchmarks in faithfulness assessment. We establish a standardized evaluation protocol incorporating attribution post-processing and cross-model consistency analysis. Empirically, BackX enables reproducible, high-discriminative evaluation across 12 state-of-the-art attribution methods, systematically exposing their causal faithfulness deficiencies. Furthermore, it inspires a novel attribution-based backdoor detection paradigm.

Developing benchmark criteria for systematic attribution assessmentEstablishing standardized evaluation setup to mitigate confounding factorsEvaluating faithfulness of attribution methods without ground truth

Latest Papers

What's happening recently
View more

This study investigates whether the "abliteration" of refusal behavior in language models selectively removes refusals without altering other decision-making characteristics. Through a weekly stock price direction prediction task—conducted in a no-refusal setting—the authors systematically compare original and abliterated models using a frozen inference pipeline, decision tendency probes, and bootstrap confidence interval analysis to assess shifts in decision bias, confidence levels, and linguistic expression. The work reveals, for the first time, consistent and reproducible off-target decisional shifts across multiple model families: abliterated models uniformly exhibit greater optimism, produce more verbose explanations, and express less uncertainty, while changes in confidence vary by model architecture. These findings challenge the assumption that refusal can be precisely excised and further demonstrate that none of the examined models possess genuine economic forecasting ability.

ablationdecision dispositionmodel surgery

Existing path-based feature attribution methods define trajectories in the input space, rendering them susceptible to path artifacts and unable to discern the semantic significance of input perturbations, which leads to unstable explanations. This work proposes Reveal-IG, a novel framework that lifts path attribution from the input space into a structured probe distribution space centered around the target sample, computing integrated gradients along distributional paths with respect to the model’s expected output. By supporting multi-scale image probes and modeling feature uncertainty in tabular data, Reveal-IG preserves attribution completeness while avoiding input-space artifacts. Experiments demonstrate that Reveal-IG produces stable, signed attributions on ImageNet classification and tabular regression tasks, significantly outperforming existing methods on sign-dependent metrics and remaining competitive on others, with synthetic diagnostics further confirming its robustness against artifacts.

attribution artifactsfeature attributioninput-space paths

This work addresses the failure of conventional single-component ablation-based attribution methods in Transformers, which overlook backup pathways due to the model’s self-repair mechanisms, often misclassifying critical components as irrelevant. To overcome this limitation, the authors propose Conditional Collaborative Ablation (CoAx), the first approach that formalizes self-repair as a conditional circuit completion task. CoAx quantifies the conditional increase in ablation effects across remaining units after removing a primary component, thereby uncovering latent second-order interactions and redundant pathways. The method operates in an unsupervised manner—requiring no labels and relying solely on model outputs—and integrates counterfactual patching, structured pruning, and cross-model transfer for precise attribution. Evaluated on the IOI task with GPT-2-small, CoAx improves AUC for backup head identification from 0.33 to 0.91, substantially outperforming baselines, and successfully generalizes across eight distinct models, enabling capability knockout and scalable pruning.

backup recoverycomponent ablationmechanistic interpretability

Hot Scholars

KW

Ke Wei

Fudan University
high dimensional signal processing and data analysisreinforcement learningnonconvex optimization
OD

Oliver Deussen

Professor of Computer Science, University of Konstanz
Computer GraphicsVisualizationModelling
KJ

Kyomin Jung

Professor, Department of Electrical and Computer Engineering, Seoul National University
Machine LearningNatural Language ProcessingSocial Network Analytics
TD

Thierry Duchesne

Professor of Statistics, Université Laval
Statistical learningGeneralized linear mixed modelsModel selectionLongitudinal data analysis
YW

Yunhai Wang

Renmin University of China
Data VisualizationHuman Data Interaction