Binary Classification with the Maximum Score Model and Linear Programming

📅 2025-07-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses binary classification under partial identification, where covariates follow a discrete distribution and model parameters are not fully identified. We propose a novel classification method grounded in Manski’s maximum score framework. Our key innovation is the first reformulation of maximum score estimation as two tractable linear programs—bypassing the computational complexity and convergence issues inherent in conventional iterative algorithms. The method retains minimal distributional assumptions (requiring neither conditional distribution continuity nor smoothness), thereby unifying computational efficiency with theoretical robustness. We establish a nontrivial finite-sample lower bound on classification accuracy and corroborate our approach via Monte Carlo simulations and empirical analysis. Results demonstrate that, relative to leading parametric and nonparametric methods, our estimator achieves superior predictive accuracy, enhanced small-sample stability, and greater robustness to model misspecification.

Technology Category

Machine Learning: Multi-class/Multi-label Learning & Extreme ClassificationSearch and Optimization: Mixed Discrete/Continuous SearchConstraint Satisfaction and Optimization: Mixed Discrete/Continuous Optimization

Application Category

Graph Algorithms and Modeling for the Web: Algorithms and analysis for incomplete, noisy, or partially observed Web-related graphsUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasets
📝 Abstract
This paper presents a computationally efficient method for binary classification using Manski's (1975,1985) maximum score model when covariates are discretely distributed and parameters are partially but not point identified. We establish conditions under which it is minimax optimal to allow for either non-classification or random classification and derive finite-sample and asymptotic lower bounds on the probability of correct classification. We also describe an extension of our method to continuous covariates. Our approach avoids the computational difficulty of maximum score estimation by reformulating the problem as two linear programs. Compared to parametric and nonparametric methods, our method balances extrapolation ability with minimal distributional assumptions. Monte Carlo simulations and empirical applications demonstrate its effectiveness and practical relevance.
Problem

Research questions and friction points this paper is trying to address.

Efficient binary classification with discrete covariates
Optimal conditions for non-classification or random classification
Linear programming reformulation for maximum score estimation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses maximum score model with linear programming
Reformulates problem as two linear programs
Balances extrapolation with minimal assumptions
💼 Related Jobs
No related jobs found.
J
Joel L. Horowitz
Department of Economics, Northwestern University, Evanston, IL 60208, USA
S
Sokbae Lee
Centre for Microdata Methods and Practice, Institute for Fiscal Studies, 7 Ridgmount Street, London, WC1E 7AE, UK; Department of Economics, Columbia University, New York, NY 10027, USA