Learning-Enabled Estimation: Tight Characterizations under Sample Selection Biases

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the unclear identifiability conditions for regression under sample selection bias, where existing methods rely on prior estimation of the selection mechanism. Drawing upon statistical learning theory, this work establishes a tight characterization of regression learnability and proposes a general estimation framework with an oracle-efficient algorithm that circumvents explicit modeling of the selection mechanism. Departing from traditional paradigms, it demonstrates that the regression function remains identifiable even when the selection mechanism is not. Furthermore, this paper derives necessary and sufficient conditions for identifiability and provides finite-sample guarantees with explicit convergence rates. The effectiveness of the proposed approach is empirically validated in practical scenarios such as auctions.
📝 Abstract
When can we learn from biased samples? We study regression when outcomes are observed only after passing through selection filters that depend on both covariates and outcomes themselves, a ubiquitous challenge spanning clinical trials with patient dropout, labor markets with self-selection, and auctions with strategic entry. Ignoring such selection yields systematically biased conclusions with real-world consequences. This challenge has a long history in econometrics and statistics, starting with Heckman's seminal two-stage model and followed by numerous generalizations. While these works provide various sufficient conditions for identification, a complete characterization of when such regression is possible has remained elusive. In this work, we provide a characterization for when regression is possible in the presence of sample selection bias. Our results establish the minimal assumptions required on the functional forms of selection processes under which regression remains possible, which are particularly relevant in modern settings where selection mechanisms are increasingly complex and opaque. As a corollary of our characterization, we show that there are settings where the regression function can be identified even when the selection filter itself cannot. This observation already goes beyond the ``estimate selection filter, then debias regression'' paradigm that is followed by virtually all existing approaches. Under natural strengthenings of our identification conditions, we also establish finite-sample estimation guarantees with explicit convergence rates and provide oracle-efficient algorithms. This yields the first general-purpose estimation method for this broad class of selection problems. Finally, we explore the implications of our results for several well-studied econometric settings with complex selection mechanisms such as auctions with entry costs and labor markets.
Problem

Research questions and friction points this paper is trying to address.

sample selection bias
regression
identification
learning from biased samples
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sample Selection Bias
Regression Identification
Oracle-Efficient Algorithms
Finite-Sample Guarantees
Heckman Model
🔎 Similar Papers
2024-07-16European Conference on Artificial IntelligenceCitations: 0