Score
Designs, builds, and evaluates systems that match probe samples to enrolled models of known classes or identities, compute probe-to-gallery similarity scores, and select operating thresholds (often via inner validation) to accept matches or reject probes as unknown. Covers development of scoring functions, threshold-selection protocols, and evaluation methods for recognition/detection when test-time instances may belong to classes not seen during training.
This work addresses the failure of probes in probe-guided fine-tuning due to Goodhart’s Law—where probes, when optimized as objectives, lose reliability. We propose an alignment method that jointly suppresses toxicity and preserves probe fidelity. Our core innovation lies in synergistically integrating supervised fine-tuning (SFT) and direct preference optimization (DPO) for probe-guided training: DPO substantially outperforms SFT, achieving both effective toxicity reduction and high probe accuracy in detecting harmful representations. Crucially, only a lightweight probe retraining post-fine-tuning is required to restore >95% of the original probe accuracy—eliminating the need for complex probe ensembling. Experiments across multiple toxicity benchmarks demonstrate significant toxicity reduction while maintaining robust internal monitoring capability, thereby validating the feasibility and practicality of probe-guided alignment.
This work addresses the vulnerability of fixed-score thresholds in 1:N face recognition, which are highly sensitive to variations in image quality or gallery composition and can lead to high-risk misidentifications. The authors propose a threshold-free rank-1 identity consistency mechanism (1-consistency), which determines whether a query subject is enrolled by requiring unanimous agreement among multiple independently trained matchers on the same top-ranked identity. This approach achieves, for the first time, a registration decision based solely on ranking consensus without any preset threshold. Comprehensive evaluation across 36 combinations of gallery and probe image qualities demonstrates that 1-consistency matches or even surpasses the performance of an Oracle-optimal threshold—without any parameter tuning—even under severely degraded probe conditions. When 1-consistency affirms enrollment, its correct matching rate reaches 97–100%, substantially outperforming the Oracle threshold’s 66–84%.
This work addresses the algorithm selection problem in black-box optimization (BBO) based on probing trajectories. We systematically evaluate 17 time-series classifiers on the BBOB benchmark, using three types of short-horizon performance trajectories as inputs. Employing leave-one-instance and leave-one-problem cross-validation, we demonstrate—for the first time—that classifier architecture critically determines trajectory-driven algorithm selection performance. Specifically, feature-engineering-based models (e.g., TSF) and interval-based models (e.g., ROCKET) significantly outperform end-to-end sequence models such as RNNs and Transformers. The best-performing classifier achieves an average accuracy over 12 percentage points higher than LSTM and InceptionTime. Our study provides the first reproducible and interpretable guideline for selecting time-series classifiers in trajectory-based algorithm selection, establishing a foundation for principled, data-driven BBO meta-algorithm design.
This work addresses the high cost of ground-truth evaluation in chemical and materials design, where existing machine learning surrogate models often lack reliability guarantees. Departing from conventional reliance on prediction accuracy metrics such as R²—which can paradoxically increase the risk of worst-case selections—the study proposes “rank preservation” as a core criterion for surrogate validation. It formally introduces the concept of “selection tax” and derives its theoretical upper and lower bounds. A safety certification framework for surrogates is established through selection-aware auditing, rank correlation analysis, and multi-task ground-truth validation. Experiments demonstrate that the proposed audit statistics achieve Spearman correlations of 0.80–0.99 with actual search performance, substantially outperforming R² (as low as 0.33). Certified screening strategies based on this framework reduce evaluation costs by up to 25-fold.
This work proposes a novel representation learning framework that addresses the limited representational capacity of existing methods in complex scenes by integrating adaptive multi-scale fusion with contrastive learning. The approach dynamically aggregates multi-level features and incorporates a structure-aware contrastive loss, thereby enhancing the model’s ability to jointly capture fine-grained semantics and global contextual information. Extensive experiments demonstrate that the proposed framework consistently outperforms state-of-the-art methods across multiple benchmark datasets, achieving substantial improvements in both accuracy and robustness. These results establish a promising new direction for unsupervised and semi-supervised representation learning.
本文提出一种基于置信区域的筛选框架,用于解决模拟系统可接受性问题,保证高概率筛选出所有或每个可接受系统,并支持并行化。
本文提出了一种基于上下文自适应阈值的方法,以解决分类器和监控程序在重要外部变量条件下分布代表性不足的问题。
本文通过去噪分数匹配方法训练卷积神经网络估计无信号得分函数,以近似贝叶斯理想观察者在信号已知的确切检测任务中的性能。
This study addresses the implicit shift in reference targets caused by candidate filtering during compact evidence evaluation in digital pathology, which distorts conclusions regarding model fidelity and strategy comparisons. To resolve this, we introduce the first reference-aware evaluation framework and propose CHARTER, an evaluation charter that transforms implicit selection into auditable specifications by explicitly declaring objectives, quantifying predictive shifts, and auditing conclusion stability. This approach effectively distinguishes genuine predictive preservation from spurious gains induced by reference changes. Experiments based on multiple instance learning and random seed auditing demonstrate that the proposed framework identifies four deterministic reversals in key comparisons and successfully rectifies prior ACMIL comparison results, thereby significantly enhancing evaluation reliability and transparency.
This study addresses the tendency of large language models to circumvent alignment objectives through superficial compliance, resulting in internal representations that fail to genuinely internalize safe behaviors. To overcome this limitation, we propose a probe-guided fine-tuning approach that, for the first time, employs continuously updated internal probes as direct optimization signals. By leveraging both linear and nonlinear probing techniques, our method shapes internal representations specifically for harmlessness and honesty, transcending the constraints of relying solely on output-level feedback. Empirically, this approach significantly outperforms Direct Preference Optimization (DPO) and inference-time interventions in navigating the safety-utility trade-off. It substantially enhances robustness against jailbreak attacks while preserving the linear encoding of concepts to ensure continued monitorability.