🤖 AI Summary
In high-stakes domains such as healthcare and criminal justice, it is often infeasible to re-conduct randomized controlled trials (RCTs) after updating machine learning models, rendering the causal effects of these updates on downstream outcomes—such as patient survival or recidivism rates—difficult to assess. This work proposes a novel partial identification approach that leverages historical RCT data and fine-grained relationships between prediction accuracy and downstream outcomes. Under two monotonicity assumptions—individual-level “counterfactual correctness” (i.e., correct predictions never lead to worse outcomes) and a trust relationship between subgroup predictive performance and outcomes—the method constructs tight bounds on the causal effect of the updated model. Simulations demonstrate that this approach yields more informative causal effect estimates compared to existing techniques.
📝 Abstract
Predictive machine learning (ML) models are increasingly used to aid human decision-makers across various high-risk domains such as healthcare and criminal justice. There is a growing recognition of the need to evaluate the causal impact of deploying these systems on downstream outcomes, such as patient survival or crime recidivism. Randomized control trials (RCTs) can provide high-quality evidence on the impact of a deployed model, but they run into a challenge: it is often infeasible to run repeated trials when models are updated or retrained to improve predictive performance. In this work, we present a partial-identification approach to using prior RCT data to construct bounds on the causal effect of a new model. The core innovation in our approach is to leverage assumptions relating fine-grained predictive accuracy to downstream outcomes. We do so via two monotonicity assumptions: first, on individual-level `counterfactual correctness' (all else being equal, a correct prediction leads to non-inferior outcomes); and second, on the relation between subgroup predictive performance and outcomes, interpretable as an assumption regarding trust in model outputs. We demonstrate our method with a simulation study, illustrating how incorporating this information can lead to more informative bounds compared to prior work.