Score
Designs, derives, and analyzes decision thresholds and thresholding strategies for detectors and classifiers, producing closed-form or data-driven threshold expressions and SNR- or sample-adaptive threshold rules. Work includes deriving analytic threshold conditions, optimizing thresholds to minimize detection or classification error, estimating thresholds from observed data, selecting operating points without exhaustive search, and performing threshold sensitivity and validation analyses.
To address low detection sensitivity and frequent omission of rare classes in heterogeneous populations, this paper proposes a covariate-dependent adaptive thresholding method. The method establishes, for the first time, a theoretical framework for the optimal adaptive threshold function, derives its asymptotic properties via nonparametric estimation, and proves a central limit theorem to support statistical inference under conditional mean and variance assumptions. It achieves precise rare-event identification through standardized statistic monitoring, proportional rule modeling, nonparametric function estimation, and bootstrap-based uncertainty quantification. Simulation and empirical studies demonstrate that, compared with conventional fixed-threshold approaches, the proposed method significantly improves alarm sensitivity while maintaining strong false-alarm control and robust inferential performance. This work provides a novel paradigm for dynamic monitoring and screening in imbalanced data settings.
Stability selection relies on manually specified stability thresholds, rendering variable selection results highly sensitive and leading to uncontrolled false discovery rates (FDR). To address this, we propose the Exclusion Automatic Threshold Selection (EATS) algorithm—a fully data-adaptive method that determines the stability threshold automatically, without prior assumptions or cross-validation. Grounded in the theoretically motivated Adaptive Threshold Selection (ATS) principle, EATS ensures statistical robustness and interpretability. The method integrates resampling-based statistics, selection probability modeling, and an exclusion-driven threshold search strategy. Extensive simulations across multiple algorithms and diverse scenarios demonstrate that EATS significantly improves selection consistency and achieves superior FDR control compared to all fixed-threshold alternatives. Moreover, EATS is plug-and-play—requiring no user tuning—and readily applicable to existing stability selection frameworks.
This paper investigates the optimal individual behavior classification problem under outcome performativity—where classifier outputs causally influence the behavior of classified agents. We formalize the decision-making process via a strategic response model and conduct rigorous theoretical analysis. Our main contribution is the first formal proof that the optimal classifier must be either a standard threshold rule or a *negative-threshold rule*, which assigns positive outcomes to agents with *lower* predicted propensity—a counterintuitive structure that strictly dominates conventional thresholding under performative settings. We establish a structural optimality theorem, constructively characterize the mechanism by which negative-threshold rules achieve superior accuracy, and generalize the result to arbitrary objective functions, including weighted misclassification losses. This finding challenges the intuitive “high-propensity-first” allocation principle and provides a provably optimal theoretical foundation and actionable design guidelines for deploying classifiers in performative environments.
This study addresses the critical issue that existing selective prediction methods in signal domains—such as anomalous sound detection and AI-generated image forensics—often yield a false sense of security due to the use of uncalibrated thresholds, resulting in actual error rates that substantially exceed users’ prescribed risk budgets. The work presents the first systematic audit of four distribution-free calibration rules (NAIVE, Hoeffding, Clopper–Pearson, and Betting) regarding their risk control performance on both real and synthetic data. Findings reveal that NAIVE exceeds the risk budget in 49–73% of experiments; Clopper–Pearson and Betting achieve zero violations under exchangeability but suffer 9–30% violation rates when deployed in grouped settings where exchangeability fails. Group-wise thresholding restores valid risk control at the cost of reduced coverage. The study underscores the pivotal role of tight confidence bounds for effective coverage and identifies uncalibrated thresholds as the root cause of risk miscontrol.
This paper addresses the challenge of analytically constructing the optimal ROC curve in binary hypothesis testing when prior distributions are unknown or intractable. We propose the first maximum likelihood estimator for the ROC curve based on observed likelihood ratio samples (MLE-ROC). Unlike conventional approaches, MLE-ROC operates directly on likelihood ratio samples without requiring knowledge of the underlying data distributions and exhibits strong convergence under the Lévy metric. We establish theoretical consistency of its AUC estimator and demonstrate—via simulations—that it significantly outperforms the empirical ROC estimator, especially in small-sample and highly imbalanced settings with sparse negative instances, reducing estimation error by over 40%. The key contribution is the first formal parameterization of the ROC curve in the likelihood ratio domain within a maximum likelihood estimation framework, accompanied by rigorous asymptotic statistical guarantees.
This work proposes a novel representation learning framework that addresses the limited representational capacity of existing methods in complex scenes by integrating adaptive multi-scale fusion with contrastive learning. The approach dynamically aggregates multi-level features and incorporates a structure-aware contrastive loss, thereby enhancing the model’s ability to jointly capture fine-grained semantics and global contextual information. Extensive experiments demonstrate that the proposed framework consistently outperforms state-of-the-art methods across multiple benchmark datasets, achieving substantial improvements in both accuracy and robustness. These results establish a promising new direction for unsupervised and semi-supervised representation learning.
This work addresses the challenge of constructing real-time review queues for risk-scoring streams in financial crime investigations by proposing a label-free, adaptive thresholding mechanism. The approach leverages online adaptive kernel density estimation (KDE), dynamically satisfying queue capacity constraints through tail-mass curves and identifying stable thresholds via persistent density minima “snapshots” detected across multiple bandwidths. Integrated with sliding windows, exponential forgetting, and priority-based queue management, the system supports multi-queue routing and real-time processing. Experimental results demonstrate that the method strictly adheres to capacity limits across synthetic, concept-drifting, and multimodal data streams, significantly reduces threshold jitter, and achieves per-event update complexity of O(G) with constant memory usage.
The area under the ROC curve (AUC) is commonly interpreted as the probability that a classifier ranks a randomly chosen positive instance higher than a negative one; however, this interpretation relies on specific assumptions whose violation can introduce bias that has not been rigorously characterized. This work systematically reviews the relevant literature and, drawing on probabilistic and statistical methods, provides the first rigorous proof of the conditions under which this probabilistic interpretation holds. Furthermore, when these assumptions are violated, the study derives a computable upper bound on the resulting deviation. By establishing a solid theoretical foundation for the probabilistic interpretation of AUC and offering explicit error bounds, this research significantly enhances the reliability and practical applicability of ROC analysis.