Which Influence Are We Estimating? The Role of Counterfactual Specifications in Data Attribution

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the long-standing misconception that estimator ranking inconsistencies in data attribution stem from approximation errors, revealing instead that they originate from counterfactual norm mismatches. By formalizing influence as a counterfactual estimator, this work establishes norm analysis as a necessary prerequisite for comparing estimators. It derives local decompositions to analytically characterize signal interaction mechanisms, validated through linearized approximations and controlled experiments. The primary contribution is the first demonstration that behavioral proxy selection critically impacts attribution quality, proving that differing norms directly induce ranking discrepancies. Furthermore, the proposed behavior-aligned norm successfully identifies target samples overlooked by default methods, substantially improving attribution accuracy.
📝 Abstract
Estimating the influence of training examples on model behavior is essential for data debugging, valuation, and attribution. Existing influence estimators often produce incompatible rankings, which are commonly ascribed to approximation error. We argue that a more fundamental source of disagreement is specification mismatch: influence depends on the behavior being attributed, the intervention applied to each training example, and the counterfactual training process that maps the intervention to a model response. These choices are especially important when the target behavior requires a tractable surrogate, such as query loss, a logit, or a margin. We formalize influence as a counterfactual estimand, distinguish specification mismatch across estimands from approximation error in estimating a fixed estimand, and organize representative estimators by their implied specifications. We further derive a local decomposition that exposes how behavior signals, training signals, and counterfactual parameter responses interact. Controlled experiments show that exact estimands under different specifications can induce different rankings, whereas approximation error grows as perturbations move farther from their linearization points. Experiments on noisy label detection and LLM attribution show that specification choices significantly affect attribution quality, especially for the choice of behavior surrogate. Behavior-aligned specifications can identify target-specific training examples obscured by default loss-based or similarity-based specifications. These results establish specification analysis as a necessary first step for interpreting and comparing data influence estimators.
Problem

Research questions and friction points this paper is trying to address.

data attribution
influence estimation
counterfactual specification
specification mismatch
training data influence
Innovation

Methods, ideas, or system contributions that make the work stand out.

Data Attribution
Counterfactual Estimand
Specification Mismatch
Influence Estimation
Behavior Surrogate
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.