Score
Choosing and tuning decision thresholds or subset-selection rules (for membership detection, abstention, or filtering) that satisfy accuracy or cost constraints, often via adaptive or uncertainty-aware procedures to balance errors and remeasurement costs.
To address low detection sensitivity and frequent omission of rare classes in heterogeneous populations, this paper proposes a covariate-dependent adaptive thresholding method. The method establishes, for the first time, a theoretical framework for the optimal adaptive threshold function, derives its asymptotic properties via nonparametric estimation, and proves a central limit theorem to support statistical inference under conditional mean and variance assumptions. It achieves precise rare-event identification through standardized statistic monitoring, proportional rule modeling, nonparametric function estimation, and bootstrap-based uncertainty quantification. Simulation and empirical studies demonstrate that, compared with conventional fixed-threshold approaches, the proposed method significantly improves alarm sensitivity while maintaining strong false-alarm control and robust inferential performance. This work provides a novel paradigm for dynamic monitoring and screening in imbalanced data settings.
This paper addresses selective classification in high-stakes settings, aiming to minimize the rejection (indecision) rate under a user-specified misclassification constraint—potentially stricter than the Bayes optimal error. We propose a threshold-adaptive framework grounded in statistical learning theory and risk-controlling optimization. By constructing tight confidence sets and dynamically adjusting decision boundaries, we establish, for the first time, theoretical guarantees on achieving the optimal rejection rate under stringent misclassification constraints. Our work challenges the conventional belief that the Bayes error is an insurmountable lower bound, and instead derives the fundamental trade-off between misclassification rate and rejection cost. Experiments demonstrate substantial reductions in misclassification—approaching zero—on hard classification tasks, while incurring only negligible rejection rates; the gain in misclassification reduction far outweighs the cost introduced by rejection.
To address fairness violations in classification arising from imbalanced error rates across protected groups, this paper proposes the Fairness-Adjusted Selection Inference (FASI) framework. FASI introduces False Selection Rate (FSR) control—a concept previously unexplored in fair classification—by provably transforming black-box model outputs into R-values, thereby guaranteeing finite-sample upper bounds on group-wise FSR and achieving statistical parity. The method integrates selection inference, R-value construction, and post-hoc optimization without requiring modifications to the underlying classifier. Experiments on synthetic and real-world datasets demonstrate that FASI substantially reduces inter-group error-rate disparities; FSR control error remains consistently below the prespecified threshold. FASI thus delivers rigorous statistical guarantees, computational efficiency, and strong fairness assurance.
This paper addresses the trade-off between robustness and efficiency under model misspecification, proposing an adaptive estimation framework that does not require a pre-specified upper bound on bias. The core challenge is to construct an estimator whose worst-case risk—relative to an oracle knowing the true bias bound—is minimized. Methodologically, we formulate an adaptive shrinkage estimator via weighted convex minimax optimization, calibrated against the oracle risk, and develop a lookup-table-based fast algorithm. Theoretically, our approach departs from conventional hypothesis-testing paradigms and achieves, for the first time, direct adaptation to the degree of misspecification. Empirically, the method substantially improves estimation accuracy and robustness across multiple canonical studies, offering both strong theoretical guarantees and practical computational efficiency.
In high-dimensional personalized treatment strategy estimation, standard post-variable-selection statistical inference fails due to selection-induced bias. Method: This paper introduces Universal Post-Selection Inference (UPoSI) into the robust Q-learning framework, proposing a selection-mechanism-agnostic universal post-selection inference method. The approach uniformly improves confidence interval construction and is theoretically shown to be asymptotically valid in multi-stage decision settings, guaranteeing nominal Type-I error control for hypothesis testing and exact coverage probability for confidence intervals. Results: Monte Carlo simulations demonstrate that the proposed method substantially improves coverage accuracy and statistical power compared to selective inference, while remaining compatible with diverse data-driven variable selection procedures. It thus provides a generalizable and verifiable foundation for statistical inference in robust Q-learning.
This study addresses the critical issue that existing selective prediction methods in signal domains—such as anomalous sound detection and AI-generated image forensics—often yield a false sense of security due to the use of uncalibrated thresholds, resulting in actual error rates that substantially exceed users’ prescribed risk budgets. The work presents the first systematic audit of four distribution-free calibration rules (NAIVE, Hoeffding, Clopper–Pearson, and Betting) regarding their risk control performance on both real and synthetic data. Findings reveal that NAIVE exceeds the risk budget in 49–73% of experiments; Clopper–Pearson and Betting achieve zero violations under exchangeability but suffer 9–30% violation rates when deployed in grouped settings where exchangeability fails. Group-wise thresholding restores valid risk control at the cost of reduced coverage. The study underscores the pivotal role of tight confidence bounds for effective coverage and identifies uncalibrated thresholds as the root cause of risk miscontrol.
This work addresses the challenge of constructing real-time review queues for risk-scoring streams in financial crime investigations by proposing a label-free, adaptive thresholding mechanism. The approach leverages online adaptive kernel density estimation (KDE), dynamically satisfying queue capacity constraints through tail-mass curves and identifying stable thresholds via persistent density minima “snapshots” detected across multiple bandwidths. Integrated with sliding windows, exponential forgetting, and priority-based queue management, the system supports multi-queue routing and real-time processing. Experimental results demonstrate that the method strictly adheres to capacity limits across synthetic, concept-drifting, and multimodal data streams, significantly reduces threshold jitter, and achieves per-event update complexity of O(G) with constant memory usage.
This study addresses the feasibility determination problem under subjective probability constraints within a finite set of alternative systems. The authors propose a statistical inference method that operates directly on Bernoulli simulation outputs, uniquely integrating multi-threshold subjective constraints with Bernoulli observations without relying on normal approximations. To handle extreme scenarios—such as when all systems are feasible or none are—the method incorporates two heuristic strategies that dynamically adjust thresholds during execution. The resulting batch-mean-independent testing algorithm maintains rigorous statistical validity while significantly outperforming existing approaches designed for normally distributed data. Empirical experiments demonstrate the method’s computational efficiency and robust adaptability across diverse problem settings.
This study addresses the challenge of simultaneously ensuring model safety, efficiency, and fairness under stringent resource constraints while complying with anti-discrimination regulations that prohibit group-dependent decision rules. The authors propose a model-agnostic post-processing framework that enforces a single global decision threshold to guarantee legal compliance and jointly optimizes multiple objectives through a parameterized ethical loss function and bounded decision rules. Theoretical analysis reveals local monotonicity of the deployed threshold with respect to ethical weights and identifies critical capacity intervals. Empirical results demonstrate that in over 80% of configurations, resource constraints dominate threshold selection; even under severe capacity limits (25%), the framework achieves high-risk identification recall rates of 0.409–0.702, substantially outperforming conventional unconstrained fairness approaches.
This work addresses the selection bias inherent in selective inference when inference is performed only on high-informativeness prediction sets, which can compromise false coverage rate (FCR) control. Under an idealized setting, the authors derive an oracle strategy that maximizes statistical power while maintaining FCR guarantees. They then develop a finite-sample calibration mechanism grounded in conformal prediction and probability calibration to adapt this optimal strategy to practical settings. The resulting procedure rigorously controls FCR while achieving substantially higher power than existing methods. Empirical evaluations on both synthetic and real-world classification datasets demonstrate consistent superiority over current approaches in terms of statistical efficiency without sacrificing FCR control.