Performative Prediction with Selective Labels

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the convergence failures and estimation biases induced by retraining mechanisms for performative prediction under selective labels. To mitigate these issues, this work proposes a confidence-interval-based robust optimization framework that integrates worst-case optimization with repeated risk minimization (RRM). By leveraging historical data to dynamically tighten confidence intervals, the approach approximates stable solutions, while introducing conditional distribution sensitivity analysis to correct conventional retraining biases. Experiments on credit fairness demonstrate that, even when only partial labels are observed, the proposed method achieves performance comparable to full-label baselines, effectively ensuring model stability and predictive reliability.
📝 Abstract
Many social applications of machine learning exhibit performative effects: population behavior changes in response to deployed models. Performative prediction studies this interaction through a distribution map that relates each model to the population distribution it induces. One of the main results in this framework showed that repeated risk minimization (RRM), which updates models by retraining on the most recent data, can converge to a stable model that minimizes risk on its own induced distribution. However, existing analyses typically assume access to the complete distributions of features and labels after model deployment, ignoring the possibility of selective labels: observing labels only for the accepted subset of the population. In this work, we formalize performative prediction with selective labels and show that retraining only on observed data can misguide the retraining procedure and undermine the guarantees of convergence to a stable solution. We then propose a worst-case objective based on knowledge of a confidence interval on the probability of a positive label. Applying RRM to this objective permits us to remain within a bounded distance to the true stable point. Under a sensitivity assumption on the conditional label distribution, we further show how previously accepted data can tighten these confidence intervals over time. Experiments in a lending application with fairness regularization show that our robust optimization approach closely matches the performance of RRM with complete label access.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Performative Prediction
Selective Labels
Repeated Risk Minimization
Robust Optimization
Confidence Intervals