Score
Designs, builds, and evaluates detectors and evaluation pipelines that identify and quantify false positive errors, including diagnosing decision-boundary failures and proposing boundary refinements; measures false-positive rates under distribution shift and model scaling, assesses lie-detector robustness and scalable oversight mechanisms, and quantifies operating-point and human-labeler cost tradeoffs to identify impractical detector settings.
This work addresses the challenge of deceptive behavior in large language models during preference learning, where human annotation is costly. It presents the first application of the scalable oversight framework SOLiD to an ultra-large-scale model (405B parameters), demonstrating its effectiveness in more realistic and diverse preference learning settings. By integrating a high-precision lie detector achieving 99% true positive rate to filter suspicious responses, the approach substantially reduces the need for human review and eliminates human annotations entirely during fine-tuning, while limiting undetected deception to 14%. The study also reveals that the method is sensitive to distributional shifts between training and preference data, which can lead to elevated false positive rates.
In industrial quality inspection, anomaly detection suffers from poor robustness due to high noise levels and sparse defective samples. To address this, we propose Iterative Refinement of Pseudo-labels (IRP), a self-supervised method that alternately evaluates sample credibility and removes misleading instances under feature-space consistency constraints—effectively purifying the training set dynamically without human annotations and generating high-fidelity self-supervised signals. IRP introduces the novel paradigm of “iterative data refinement,” significantly enhancing model robustness against label noise and cross-domain generalization capability. Evaluated on KSDD2 and MVTec AD benchmarks, IRP consistently outperforms existing unsupervised and self-supervised methods. Notably, under high-noise conditions, it achieves substantial improvements in detection accuracy and reduces false positive rates by over 25%.
This study addresses the challenge large language models face in detecting cross-paragraph structural inconsistencies during multi-agent collaborative long-document generation. By fixing document content, defect types, and evaluation protocols, the authors systematically evaluate ten prominent models under both single-agent and multi-agent settings. Leveraging signal detection theory decomposition, controlled experiments, private record reconstruction, and automated scoring, they find that all models exhibit a performance drop exceeding two-thirds in multi-agent coordination scenarios. Only one developer-provided model shows a significant shift in reporting criteria (p<0.001), yet its confidence scores fail to reflect cross-segment defects. This work is the first to reveal the “detection cliff” phenomenon and introduces a reproducible benchmark framework for future research.
This work addresses the performance gap between academic benchmarks and real-world deployment in unsupervised anomaly detection, where existing methods often exhibit instability, sensitivity to preprocessing, and inconsistent behavior in industrial settings. The authors conduct a systematic evaluation of 19 models on BowTie, a complex manufacturing dataset, revealing significant discrepancies between benchmark results and practical efficacy. To bridge this gap, they propose a human-in-the-loop unified detection framework that integrates SAM-generated refined candidate regions, heatmap-guided inspection, mask-based evaluation, and interactive verification. This framework enables quality inspectors to efficiently confirm defects, refine boundaries, and trace historical cases. Preliminary deployment demonstrates that the system substantially enhances both reliability and efficiency in industrial visual inspection.
The core challenge in data quality monitoring lies in error provenance—specifically, identifying the underlying mechanisms that generate errors—a problem largely overlooked by existing work, which seldom models such mechanisms explicitly. This paper focuses on errors arising from intrinsic dependencies within data and proposes MechDetect, the first method to systematically extend missing-data mechanism detection to diverse error types—including outliers, inconsistencies, and format violations. Leveraging joint statistical modeling and supervised learning, MechDetect simultaneously models tabular data and their error masks to automatically determine whether observed errors stem from inherent characteristics of the original data. Extensive experiments across multiple benchmark datasets demonstrate that MechDetect significantly outperforms state-of-the-art baselines in accurately diagnosing error-generation mechanisms. By providing mechanistic interpretability, it establishes a theoretical foundation and practical framework for explainable data repair.
Current AI evaluation frameworks overemphasize output correctness while neglecting the resource costs required to verify errors in real-world deployment, allowing high accuracy metrics to mask substantial verification burdens. This work introduces verification-cost errors (VCEs)—errors that cannot be detected by a specified proportion of validators within a given verification budget—thereby shifting the paradigm from defining errors solely by output properties to centering on their detectability during verification. Through an operational definition, verification budget modeling, and user studies, we empirically demonstrate in code generation and multimodal document understanding tasks that high benchmark accuracy can coexist with significant verification effort, underscoring that correctness alone is insufficient to reflect system reliability in practical settings.
This study addresses the challenge of alarm fatigue in continuous model monitoring caused by high false positive rates of existing drift detectors, which undermines monitoring reliability. It presents the first systematic evaluation of the cumulative false positive behavior of five widely used methods—Population Stability Index (PSI), Kolmogorov–Smirnov (KS) test, Maximum Mean Discrepancy (MMD), Least-Squares Density Difference (LSDD), and adversarial validation—under continuous monitoring settings, incorporating Bonferroni correction for multiple hypothesis testing. The empirical analysis reveals that PSI exhibits markedly improved stability when sample sizes exceed 200, whereas KS, MMD, and LSDD demonstrate greater reliability with smaller batch sizes. While Bonferroni correction effectively suppresses false positives, it concurrently reduces detection sensitivity. These findings offer practical guidance for selecting batch sizes and calibrating detectors in real-world deployments, balancing robustness and responsiveness.