Score
Designs and implements automated procedures that select numeric decision thresholds for model scores or probabilistic outputs, producing reproducible cutoff rules derived from validation metrics, calibration, or statistical criteria. Builds algorithms to tune thresholds offline or online, combine multiple selection rules, and optimize trade-offs (e.g., precision–recall, false‑alarm rate) while adapting thresholds to changing operating conditions or data distributions.
Stability selection relies on manually specified stability thresholds, rendering variable selection results highly sensitive and leading to uncontrolled false discovery rates (FDR). To address this, we propose the Exclusion Automatic Threshold Selection (EATS) algorithm—a fully data-adaptive method that determines the stability threshold automatically, without prior assumptions or cross-validation. Grounded in the theoretically motivated Adaptive Threshold Selection (ATS) principle, EATS ensures statistical robustness and interpretability. The method integrates resampling-based statistics, selection probability modeling, and an exclusion-driven threshold search strategy. Extensive simulations across multiple algorithms and diverse scenarios demonstrate that EATS significantly improves selection consistency and achieves superior FDR control compared to all fixed-threshold alternatives. Moreover, EATS is plug-and-play—requiring no user tuning—and readily applicable to existing stability selection frameworks.
本文解决了机器学习系统中因数据聚类导致阈值设定不准确的问题,通过提出一种新的有效样本量计算方法来修正阈值设定。
To address low detection sensitivity and frequent omission of rare classes in heterogeneous populations, this paper proposes a covariate-dependent adaptive thresholding method. The method establishes, for the first time, a theoretical framework for the optimal adaptive threshold function, derives its asymptotic properties via nonparametric estimation, and proves a central limit theorem to support statistical inference under conditional mean and variance assumptions. It achieves precise rare-event identification through standardized statistic monitoring, proportional rule modeling, nonparametric function estimation, and bootstrap-based uncertainty quantification. Simulation and empirical studies demonstrate that, compared with conventional fixed-threshold approaches, the proposed method significantly improves alarm sensitivity while maintaining strong false-alarm control and robust inferential performance. This work provides a novel paradigm for dynamic monitoring and screening in imbalanced data settings.
本文提出了一种基于上下文自适应阈值的方法,以解决分类器和监控程序在重要外部变量条件下分布代表性不足的问题。
This work proposes a novel representation learning framework that addresses the limited representational capacity of existing methods in complex scenes by integrating adaptive multi-scale fusion with contrastive learning. The approach dynamically aggregates multi-level features and incorporates a structure-aware contrastive loss, thereby enhancing the model’s ability to jointly capture fine-grained semantics and global contextual information. Extensive experiments demonstrate that the proposed framework consistently outperforms state-of-the-art methods across multiple benchmark datasets, achieving substantial improvements in both accuracy and robustness. These results establish a promising new direction for unsupervised and semi-supervised representation learning.
This study addresses the critical issue that existing selective prediction methods in signal domains—such as anomalous sound detection and AI-generated image forensics—often yield a false sense of security due to the use of uncalibrated thresholds, resulting in actual error rates that substantially exceed users’ prescribed risk budgets. The work presents the first systematic audit of four distribution-free calibration rules (NAIVE, Hoeffding, Clopper–Pearson, and Betting) regarding their risk control performance on both real and synthetic data. Findings reveal that NAIVE exceeds the risk budget in 49–73% of experiments; Clopper–Pearson and Betting achieve zero violations under exchangeability but suffer 9–30% violation rates when deployed in grouped settings where exchangeability fails. Group-wise thresholding restores valid risk control at the cost of reduced coverage. The study underscores the pivotal role of tight confidence bounds for effective coverage and identifies uncalibrated thresholds as the root cause of risk miscontrol.
This work addresses the limitation of existing safety-critical systems, which typically evaluate only predictive accuracy while lacking rigorous validation of the overall calibration of predicted probability distributions. To bridge this gap, the authors propose a modular calibration testing framework that decouples the calibration process into four interchangeable components: data model, scoring rule, hypothesis formulation, and statistical test procedure. Built upon formal statistical hypothesis testing, the framework provides a single accept/reject decision for the entire predictive distribution. Crucially, it rejects only overly confident predictions while tolerating reasonable deviations, thereby balancing practicality with flexibility. Empirical evaluations on weather forecasting and robotic pose estimation tasks demonstrate that the framework effectively supports reliable deployment in safety-critical applications.
This study addresses the unreliability of conclusions regarding class-imbalance methods derived from single-dataset evaluations. Employing a leakage-free nested cross-validation protocol across 45 binary classification tasks, we conducted large-scale experiments to reassess these techniques. Results reveal that threshold tuning benefits exhibit non-monotonic variation with imbalance ratios and refute the hypothesis that calibration error predicts tuning gains. Furthermore, Random Forest combined with SMOTE demonstrated significant effectiveness across multiple tasks, highlighting the limitations of findings based solely on single fraud datasets. This work systematically clarifies the true utility and applicability boundaries of resampling and threshold tuning, providing robust empirical evidence for the reliable evaluation of imbalanced classification methods.
This work proposes AutoSI, a novel framework that automates selective inference for any algorithm whose selection event can be expressed as a rational function of the data, eliminating the need for manual derivation by experts. By modeling selection events through rational functions and integrating automatic symbolic computation with exact finite-sample p-value calculation, AutoSI overcomes the limitations of existing methods, which are typically confined to linear or quadratic inequalities. Empirical evaluations across three feature selection tasks—including Lasso tuned via cross-validated R²—demonstrate that AutoSI rigorously controls Type I error while maintaining high statistical power, thereby offering a general, scalable solution for post-selection inference without human intervention.
This study addresses the challenge of simultaneously controlling both types of error in binary classification tasks, where overlapping class distributions inherently hinder such control. Within the Neyman-Pearson framework, this work introduces a rejection mechanism and proposes a model-agnostic joint calibration strategy to achieve selective dual error rate control. Furthermore, by leveraging martingale theory to precisely compute finite-sample crossing probabilities, the method ensures strict statistical validity without requiring multiple testing corrections. This approach overcomes the limitations of conventional threshold search procedures by guaranteeing that both error types strictly satisfy their predefined bounds. Its effectiveness is empirically validated in high-stakes applications, including recidivism prediction and credit default assessment, thereby providing reliable statistical safeguards for critical decision-making scenarios.