🤖 AI Summary
This study addresses the problem of security monitors generating excessive false alarms due to conservative policies, which wastes auditing resources and undermines user trust. We formulate false alarm auditing as a Positive-Unlabeled (PU) ranking task and propose a two-stage learning framework. Key innovations include introducing trust-aware PU supervision and reliability-gated ranking distillation to overcome monitor-induced selection bias, alongside integrating hierarchical structure refinement with multi-reference model consensus distillation to enable accurate false alarm identification without alarm labels. Experimental results demonstrate that the proposed method achieves a macro-averaged AUPRC of 0.6444, outperforming baselines by 5.27–16.98 percentage points, and recovers 33.3% more false alarms under a 5% auditing budget.
📝 Abstract
Safety monitors help safeguard language-model agents interacting with external tools and environments, but conservative monitoring can generate many false alarms, consuming extensive review resources and weakening trust in alerts. Because false and genuine alarms often remain interleaved in native monitor scores, obtaining a reliable cutoff still requires substantial manual verification. In practice, a small set of verified-safe non-alarmed trajectories may be available while alarms remain unlabeled, naturally casting false-alarm auditing as a positive-unlabeled (PU) ranking problem. The key challenge is monitor-induced selection, since observed safe references are accepted by the monitor, while the hidden safe alarms of interest are precisely those it incorrectly flags, making the observed positives poorly representative of the positives to be recovered. To address this challenge, we propose a two-stage framework in which Trust-aware PU Supervision adapts safe references toward the alarm domain and protects plausible false alarms from excessive negative pressure, while Reliability-gated Rank Distillation consolidates consistent ordering preferences from multiple PU reference models into a single student. Consensus-guided Structural Refinement then improves the student ranking using hierarchical safe-reference support, alarm relations, and predicted reference consensus. The framework requires no alarm safety labels for fitting and leaves the underlying monitor unchanged. Across mainstream safety monitors, our method achieves a macro AUPRC of $0.6444$, outperforming eight evaluated PU baselines by 5.27--16.98 absolute percentage points; compared with PULDA, the strongest evaluated PU baseline, it recovers 33.3% more false alarms at a 5% review budget.