Data-Poisoning Audits for Causal Effect Estimation

πŸ“… 2026-07-21
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the vulnerability of cross-site causal analysis to append-only data poisoning attacks, wherein adversaries inject plausible yet carefully crafted records to distort treatment effect estimates. The authors propose a poisoning audit framework tailored to augmented inverse probability weighting (AIPW) estimators, which precisely quantifies the worst-case causal effect bias under constraints on record plausibility, poisoning budget, and source capacity. Key contributions include a greedy scanning algorithm that efficiently computes the worst-case bias for any finite budget and sample size, and a novel Total Influence Score that unifies the direct and indirect impacts of individual records on both the propensity score and outcome models. Notably, this score yields the first conservative finite-budget bound for fully re-fitted estimators. Experiments demonstrate that the framework accurately predicts bias and that even minimal poisoning budgets can substantially compromise causal inference across multiple real-world and public datasets.
πŸ“ Abstract
Observational causal analyses increasingly pool records across sites, vendors, and collection systems, creating vulnerability to append-only attacks in which plausible records are strategically selected to alter a reported treatment effect. We develop a data-poisoning audit for augmented inverse-probability-weighted estimation. The analyst specifies a finite catalog of feasible records, an append budget, and nested source capacities, and the adversary selects a feasible subset to maximize movement in a prespecified direction. With preprocessing and nuisance fits held fixed, we propose a greedy scan that computes the exact finite-sample worst-case movement at every append budget. To account for nuisance refitting, we go on to derive a total-influence score combining each record's direct contribution with its effect through the propensity and outcome models. We further obtain a conservative finite-budget bound for the fully refitted estimate. Extensive simulations validate the exact result and show that total influence improves local refit prediction, while multisite and public-data analyses demonstrate material sensitivity at small append budgets. By translating adversarial data-composition risk into movement curves and critical budgets, the framework supports more reliable causal reporting and the design of source-level safeguards.
Problem

Research questions and friction points this paper is trying to address.

data poisoning
causal effect estimation
observational studies
adversarial attacks
append-only attacks
Innovation

Methods, ideas, or system contributions that make the work stand out.

data poisoning
causal inference
adversarial robustness
influence function
augmented inverse probability weighting
πŸ”Ž Similar Papers
No similar papers found.