🤖 AI Summary
This paper addresses binary classification under partial identification, where covariates follow a discrete distribution and model parameters are not fully identified. We propose a novel classification method grounded in Manski’s maximum score framework. Our key innovation is the first reformulation of maximum score estimation as two tractable linear programs—bypassing the computational complexity and convergence issues inherent in conventional iterative algorithms. The method retains minimal distributional assumptions (requiring neither conditional distribution continuity nor smoothness), thereby unifying computational efficiency with theoretical robustness. We establish a nontrivial finite-sample lower bound on classification accuracy and corroborate our approach via Monte Carlo simulations and empirical analysis. Results demonstrate that, relative to leading parametric and nonparametric methods, our estimator achieves superior predictive accuracy, enhanced small-sample stability, and greater robustness to model misspecification.
📝 Abstract
This paper presents a computationally efficient method for binary classification using Manski's (1975,1985) maximum score model when covariates are discretely distributed and parameters are partially but not point identified. We establish conditions under which it is minimax optimal to allow for either non-classification or random classification and derive finite-sample and asymptotic lower bounds on the probability of correct classification. We also describe an extension of our method to continuous covariates. Our approach avoids the computational difficulty of maximum score estimation by reformulating the problem as two linear programs. Compared to parametric and nonparametric methods, our method balances extrapolation ability with minimal distributional assumptions. Monte Carlo simulations and empirical applications demonstrate its effectiveness and practical relevance.