Score
Designs, builds, and evaluates systems that detect precursors to adverse events and issue timely alerts; this includes selecting sensors and data sources, defining detection and forecasting algorithms and decision thresholds, specifying notification and escalation rules, and setting performance metrics such as lead time, sensitivity, false‑alarm rate, latency, and robustness. Implements and tests signal processing, anomaly detection, risk modeling, simulation and calibration, and operational workflows to ensure reliable, timely warnings with acceptable trade‑offs between early detection and false positives.
Naval systems frequently exhibit anomalous behaviors due to wear, misuse, or component failures—challenges that hinder timely detection and precise remediation. To address this, we propose a predictive-diagnostic closed-loop framework that tightly integrates the existing failure prediction system PREVENT with a newly designed responsive troubleshooting module, REACT. Methodologically, the framework synergizes multi-source time-series anomaly detection with domain-knowledge-driven fault-isolation process modeling, enabling end-to-end automation—from anomaly alerting and root-cause localization to actionable remediation recommendations. Evaluated on operational shipboard systems deployed by Fincantieri, the framework reduces mean time to fault localization by 42%, significantly improves operational response efficiency, and demonstrates strong generalizability across diverse industrial domains.
In cloud-native systems, alert rules frequently suffer from false positives and false negatives due to the absence of design-phase validation, while existing tools lack systematic support for alert testing. To address this, we propose the “Alert-as-Experiment” paradigm—the first adaptation of the observability experimentation framework OXN to early-stage alert rule validation. Our approach enables closed-loop, development-time testing and continuous calibration of alert logic via simulated execution, synthetic observation data injection, and real-world scenario replay. It supports parameter tuning and repeatable verification of alert-triggering behavior, shifting alert engineering from empirical practice toward a testable, verifiable, and systematic discipline. Empirical evaluation demonstrates significant reductions in both false positive and false negative rates in production environments, alongside improved fault response latency and system maintainability.
Large language models (LLMs) can still generate unsafe outputs in deployment, necessitating efficient real-time monitoring. This work proposes a lightweight online safety monitoring mechanism that integrates signals from an external verification model, threshold-based decision rules, and risk control theory to produce reliable alerts through calibrated thresholds. The approach features a simple architecture that avoids computationally intensive procedures yet achieves detection performance on par with state-of-the-art sequential hypothesis testing methods across mathematical reasoning and red-teaming benchmarks. By combining practical efficiency with theoretical guarantees, the proposed method offers a viable solution for real-world LLM safety monitoring.
Safety-critical cyber-physical systems (CPS), such as artificial pancreas systems (APS), face urgent challenges in real-time prediction and proactive mitigation of safety hazards caused by malicious attacks or unexpected failures. Method: We propose a tightly integrated, knowledge-guided and data-driven safety engine. It introduces, for the first time, a closed-loop framework unifying joint estimation of short- and long-term system trajectories, causal inference of latent safety hazards, and generation of optimal corrective actions—incorporating domain-specific safety constraint knowledge graphs, context-aware mitigation policy libraries, temporal deep learning models (LSTM/TCN), and optimization-based action planning. Contribution/Results: Evaluated on a real-world APS testbed and clinical datasets, our approach achieves a 92.8% hazard mitigation success rate—improving over pure rule-based or pure data-driven baselines by >76%. It guarantees zero false negatives, maintains low false positive rates, and introduces no new safety risks.
This work addresses the challenge of abrupt temporal metric anomalies in large-scale base station testing, often caused by resource allocation errors, which necessitate efficient unsupervised detection methods. The authors propose CALM, a framework that leverages nonparametric kernel density estimation combined with bootstrap-based dynamic thresholding to enable real-time anomaly detection at the individual testbed level. To mitigate alert fatigue and identify globally significant anomalies, they further introduce AggCALM, which aggregates signals across multiple testbeds. Requiring no labeled data, the approach offers high timeliness and scalability, demonstrating strong performance on both simulated and real-world base station datasets. Beyond ensuring stable operation in complex testing environments, the method is readily transferable to other system health monitoring scenarios.
This work addresses the confounding of physical process modeling capability and alarm thresholding effects in existing evaluations of cyber-physical system (CPS) anomaly detectors, which obscures the attribution of performance differences. To resolve this, the authors decouple detection into two stages—residual generation and threshold-based alarming—and propose a normalized residual energy–based evaluation metric. This metric independently quantifies a model’s ability to represent the underlying physical process without relying on specific decision rules or hyperparameter tuning. Furthermore, it connects to KL divergence to measure attack separability, training–testing stability, and model compactness. Evaluations across five detector families on the SWaT, WADI, and HAI benchmarks reveal that performance rankings are highly scenario-dependent and precisely identify failure causes—such as inadequate representation capacity, suboptimal thresholds, or weak physical manifestations of attacks.
This study addresses the critical issue that existing selective prediction methods in signal domains—such as anomalous sound detection and AI-generated image forensics—often yield a false sense of security due to the use of uncalibrated thresholds, resulting in actual error rates that substantially exceed users’ prescribed risk budgets. The work presents the first systematic audit of four distribution-free calibration rules (NAIVE, Hoeffding, Clopper–Pearson, and Betting) regarding their risk control performance on both real and synthetic data. Findings reveal that NAIVE exceeds the risk budget in 49–73% of experiments; Clopper–Pearson and Betting achieve zero violations under exchangeability but suffer 9–30% violation rates when deployed in grouped settings where exchangeability fails. Group-wise thresholding restores valid risk control at the cost of reduced coverage. The study underscores the pivotal role of tight confidence bounds for effective coverage and identifies uncalibrated thresholds as the root cause of risk miscontrol.
This study addresses the challenge that traditional anomaly detection methods struggle to identify samples near the boundary between normal and anomalous states, thereby failing to enable early fault warnings. To overcome this limitation, the paper introduces the novel concept of “near-anomalies” and proposes CANARI, an unsupervised method grounded in Christoffel function theory to model such borderline cases. By moving beyond conventional dual-threshold mechanisms, CANARI proactively identifies unlabeled samples likely to evolve into failures. Experimental results on both synthetic and real-world industrial printed circuit board in-circuit test data demonstrate that CANARI significantly outperforms existing baselines, offering a robust foundation for predictive maintenance and quality control.