Score
Designs, builds, and analyzes Neyman–Pearson–style classification and decision procedures that enforce a specified constraint on an accuracy-related quantity (e.g., type I error or other loss) in expectation while optimizing a secondary performance metric. Work includes calibrating thresholds to meet the expected target, deriving finite-sample and asymptotic guarantees, and correcting bias that arises in empirical utility maximization.
This work addresses the challenge in Neyman–Pearson (NP) classification of reliably constraining the type I error (i.e., misclassification rate of the prioritized class) under finite-sample settings, where existing methods often suffer from insufficient control. The authors identify this issue as stemming from an over-optimism bias inherent in statistical learning and propose two novel algorithms: one that enforces type I error control in expectation and another that guarantees such control with high probability. By integrating asymptotic theory, empirical risk minimization, and confidence bound construction, the proposed framework yields the first NP classifiers with rigorous theoretical guarantees and tight error control. It further incorporates class-specific accuracy estimation and uncertainty quantification. Experiments demonstrate that the methods substantially improve type I error control in finite samples and enhance screening reliability in real-world cancer detection tasks.
This work addresses the fundamental tension between fairness and economic efficiency in automated decision-making. Methodologically, it introduces the NP-EO paradigm—the first framework to rigorously incorporate Equal Opportunity (EO) constraints into the Neyman–Pearson (NP) classification framework. We formulate a joint optimization problem, derive the oracle classifier satisfying both EO and NP constraints, and design a finite-sample algorithm that, with high probability, guarantees group-level fairness (strict EO compliance) and controlled Type I error rate. Theoretically and empirically—across synthetic and real-world datasets—the algorithm achieves statistical validity (precise error control) and social validity (substantial improvement in true positive rates for disadvantaged groups), while incurring only bounded efficiency loss. The core contribution is the first exact integration of NP decision theory with group-fairness constraints, yielding a verifiable, deployable solution for high-stakes classification tasks.
This work addresses the reliability challenge of selective classification under covariate shift. We propose a likelihood-ratio-based rejection mechanism grounded in the Neyman–Pearson lemma: when the test distribution deviates from the training distribution, the model abstains from prediction upon insufficient predictive confidence. To our knowledge, this is the first systematic integration of NP-optimal hypothesis testing theory into selective classification—unifying multi-class baselines and designing a novel selection function tailored for covariate shift. Our approach synergistically combines posterior calibration with out-of-distribution detection to enhance the robustness of rejection decisions. Extensive experiments across vision, language, and vision-language modeling (VLM) tasks demonstrate significant improvements over state-of-the-art methods. Results validate that the likelihood-ratio strategy effectively enhances both predictive reliability and generalization under distributional shift, exhibiting broad applicability across modalities and architectures.
This paper addresses the optimal decision problem for composite binary hypothesis testing under the Neyman–Pearson framework: maximizing the expected value of a nonlinear function of the detection probability subject to a false-alarm probability constraint. Methodologically, it establishes a novel equivalence—under a generalized Bayesian perspective—between power functions and generalized Bayes rules, enabling the construction of a weighted likelihood ratio test with a single threshold applicable to both composite null and composite alternative hypotheses. The framework unifies treatment of average- and worst-case false-alarm constraints. By leveraging signed measure integration optimization and exponential-family structural analysis, the authors derive an explicit analytical form for the optimal threshold, yielding closed-form solutions within exponential families. Numerical experiments demonstrate substantial improvements in detection performance while rigorously satisfying the prescribed false-alarm constraints.
This paper addresses the trade-off between robustness and efficiency under model misspecification, proposing an adaptive estimation framework that does not require a pre-specified upper bound on bias. The core challenge is to construct an estimator whose worst-case risk—relative to an oracle knowing the true bias bound—is minimized. Methodologically, we formulate an adaptive shrinkage estimator via weighted convex minimax optimization, calibrated against the oracle risk, and develop a lookup-table-based fast algorithm. Theoretically, our approach departs from conventional hypothesis-testing paradigms and achieves, for the first time, direct adaptation to the degree of misspecification. Empirically, the method substantially improves estimation accuracy and robustness across multiple canonical studies, offering both strong theoretical guarantees and practical computational efficiency.
This work addresses the challenge of class-specific error rate control in multiclass Neyman–Pearson (NP) classification under label noise by proposing a novel empirical likelihood–based approach. The method models the relationship between noisy and true label distributions via an exponential tilting density ratio, and jointly estimates the true posterior probabilities and class priors through an EM algorithm combined with nonparametric inference, yielding consistent and asymptotically normal estimators. These estimates are then used to construct a classifier that satisfies pre-specified class-wise error constraints. To the best of our knowledge, this is the first study to integrate empirical likelihood into the noisy-label multiclass NP classification framework, with theoretical guarantees that the resulting classifier satisfies the NP oracle inequality. Experiments demonstrate that the proposed method achieves near-oracle performance in simulations and substantially outperforms existing approaches that ignore label noise.
This study addresses the challenge of efficiently guiding treatment allocation in a main experiment using a small-scale pilot, avoiding efficiency losses from noise or excessive conservatism. The authors propose the Conditional Minimax Regret (CMR) rule, which optimizes assignment probabilities in a two-stage design by leveraging confidence sets constructed from limited pilot data, thereby balancing robustness and adaptivity. The CMR rule preserves, with high probability, the worst-case guarantees of balanced designs while asymptotically converging to the Neyman allocation as the pilot sample size grows, achieving the minimax regret rate. The approach naturally extends to multi-arm and stratified settings. Simulations demonstrate that CMR substantially outperforms feasible Neyman allocation when the pilot is small—avoiding its severe precision loss—while recovering most of its efficiency gains in large samples.
This work addresses the potential nonexistence of a global optimum in linear ensembles of multiple binary classifiers by proposing a theoretical framework grounded in truth-table logical structuring and equivalence class partitioning, which establishes sufficient conditions for the existence of a convexified empirical risk minimizer. By introducing a multidimensional generalization of classification-calibrated loss functions and the notion of φ-frontiers, the study analyzes solution stability in relation to data quality. Under exponential (Boost) and logistic (Logit) losses, the authors derive, for the first time, explicit closed-form expressions for the optimal ensemble weights and fully characterize all solution regimes in the three-classifier setting. This approach circumvents iterative optimization, thereby substantially enhancing both the interpretability and computational efficiency of ensemble models.
This work addresses the limitations of traditional generalization analyses, which rely on the often unverifiable assumption of independent and identically distributed (i.i.d.) data and thus struggle to accurately characterize model performance on unseen data. The paper proposes a deterministic generalization analysis framework that dispenses with any prior probabilistic assumptions. By examining the sensitivity of optimization solutions to data perturbations, it decomposes the generalization error into geometric and probabilistic components, achieving their first-ever decoupling. The framework expresses generalization bounds via a variational principle, leveraging deterministic perturbation analysis and optimization sensitivity theory to capture the discrepancy between in-sample and out-of-sample performance. Error terms are evaluated through posterior statistical hypotheses, enabling the recovery of conventional high-probability or expected generalization guarantees—all without requiring distributional assumptions.