Repro Samples Method for Model-Free Inference in High-Dimensional Binary Classification

📅 2025-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Statistical inference for high-dimensional binary classification remains challenging without restrictive assumptions—such as logistic/probit regression or parameter sparsity. Method: This paper proposes a model-agnostic inferential framework grounded in the repro samples paradigm. It generates synthetic samples via a sparse-constrained generalized linear working model; crucially, valid finite-sample coverage is guaranteed even when the working model is misspecified. The framework constructs confidence sets for arbitrary linear combinations of regression coefficients and yields a candidate model set with provable coverage. Contribution/Results: It is the first method to achieve consistent inference for the model support, individual prediction probabilities, and oracle coefficients—without requiring sparsity or strong parametric assumptions. Empirically, the candidate set is compact and efficient; under correct model specification, it outperforms state-of-the-art debiased estimators. Applied to single-cell RNA-seq data, it identifies a previously unreported, statistically significant immune-related gene, revealing a novel immune response mechanism.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationReasoning under Uncertainty: Relational Probabilistic ModelsCognitive Modeling & Cognitive Systems: Conceptual Inference and Reasoning

Application Category

User Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalizationSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Representation, reconstruction, and subgraph or motif discovery in Web-related graphs
📝 Abstract
This paper presents a novel method for statistical inference in high-dimensional binary models with unspecified structure, where we leverage a (potentially misspecified) sparsity-constrained working generalized linear model (GLM) to facilitate the inference process. Our method is based on the repro samples framework, which generates artificial samples that mimic the actual data-generating process. Our inference targets include the model support, case probabilities, and the oracle regression coefficients defined in the working GLM. The proposed method has three major advantages. First, this approach is model-free, that is, it does not rely on specific model assumptions such as logistic or probit regression, nor does it require sparsity assumptions on the underlying model. Second, for model support, we construct a model candidate set for the most influential covariates that achieves guaranteed coverage under a weak signal strength assumption. Third, for oracle regression coefficients, we establish confidence sets for any group of linear combinations of regression coefficients. Simulation results demonstrate that the proposed method produces valid and small model candidate sets. It also achieves better coverage for regression coefficients than the state-of-the-art debiasing methods when the working model is the actual model that generates the sample data. Additionally, we analyze single-cell RNA-seq data on the immune response. Besides identifying genes previously proven as relevant in the literature, our method also discovers a significant gene that has not been studied before, revealing a potential new direction in understanding cellular immune response mechanisms.
Problem

Research questions and friction points this paper is trying to address.

Model-free inference for high-dimensional binary classification
Constructing valid confidence sets for regression coefficients
Identifying influential covariates with guaranteed coverage guarantees
Innovation

Methods, ideas, or system contributions that make the work stand out.

Model-free inference using repro samples framework
Guaranteed coverage for influential covariate sets
Confidence sets for linear combinations of coefficients
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
X
Xiaotian Hou
Department of Statistics, Rutgers, The State University of New Jersey, Piscataway, NJ 08854
P
Peng Wang
Department of Operations, Business Analytics, and Information Systems, University of Cincinnati, Cincinnati, OH 45221
Minge Xie
Minge Xie
Rutgers
statisticsdata science
Linjun Zhang
Linjun Zhang
Associate Professor of Statistics, Rutgers University
High-Dimensional StatisticsDeep LearningDifferential PrivacyAlgorithmic Fairness