Score
Designs and implements systems that continuously observe a running system’s outputs and verifier signals, compute safety checks or scores in real time, compare them to thresholds or policies, and trigger alerts, mitigations, or logging when unsafe conditions are detected. Focuses on low-latency monitoring, reliable signal processing, thresholding and decision logic, and integration with runtime components to ensure timely detection and response to unsafe outputs.
Large language models (LLMs) can still generate unsafe outputs in deployment, necessitating efficient real-time monitoring. This work proposes a lightweight online safety monitoring mechanism that integrates signals from an external verification model, threshold-based decision rules, and risk control theory to produce reliable alerts through calibrated thresholds. The approach features a simple architecture that avoids computationally intensive procedures yet achieves detection performance on par with state-of-the-art sequential hypothesis testing methods across mathematical reasoning and red-teaming benchmarks. By combining practical efficiency with theoretical guarantees, the proposed method offers a viable solution for real-world LLM safety monitoring.
Runtime errors caused by temporal mismatches—such as synchronization failures in cyber-physical systems (e.g., drones)—pose critical reliability challenges in asynchronous data stream monitoring. Method: We propose a novel pacing-based type system that formally models the timing behavior of a core fragment of RTLola and rigorously proves its type safety. This system integrates pacing constraints directly into typing rules—enabling precise detection of asynchronous synchronization errors that elude conventional approaches—and combines static analysis to verify fine-grained data synchronization strategies. Contribution/Results: Our implementation performs whole-program static checking of RTLola monitors, eliminating undefined behavior induced by asynchronous inputs at compile time. This significantly enhances the reliability and safety of streaming monitoring components, marking the first type system to formally incorporate pacing for asynchronous stream synchronization verification.
Existing runtime enforcement techniques struggle to handle reactive systems with complex continuous dynamics and lack effective mechanisms for intervening in hybrid behaviors. This work proposes the first framework that integrates hybrid automata into runtime enforcement, enabling coordinated discrete event editing and continuous-time monitoring to correct system behavior at any instant by suppressing, delaying, or inserting events. The paper establishes formal enforceability conditions and devises an online strategy synthesis algorithm based on reachability analysis. Evaluation on an adaptive cruise control case study demonstrates that the approach ensures safety properties even when the underlying controller is unsafe, all while incurring minimal computational overhead.
This paper addresses the problem of optimally composing multiple runtime monitors under an average cost constraint to maximize the safety intervention probability (i.e., recall) against AI misaligned outputs. The proposed method introduces a Neyman–Pearson lemma–based optimization framework that unifies monitor invocation timing, selection, and intervention decisions into a likelihood-ratio–driven sequential decision problem. Pareto-optimal solutions are identified via exhaustive search, enabling principled trade-offs between performance and computational cost. Empirical evaluation on code review tasks demonstrates that the approach significantly improves multi-monitor coordination efficiency, achieving over 100% recall improvement relative to baseline methods. These results validate both the theoretical soundness and practical efficacy of the framework in resource-constrained real-world deployment scenarios.
Safety-critical cyber-physical systems (CPS), such as artificial pancreas systems (APS), face urgent challenges in real-time prediction and proactive mitigation of safety hazards caused by malicious attacks or unexpected failures. Method: We propose a tightly integrated, knowledge-guided and data-driven safety engine. It introduces, for the first time, a closed-loop framework unifying joint estimation of short- and long-term system trajectories, causal inference of latent safety hazards, and generation of optimal corrective actions—incorporating domain-specific safety constraint knowledge graphs, context-aware mitigation policy libraries, temporal deep learning models (LSTM/TCN), and optimization-based action planning. Contribution/Results: Evaluated on a real-world APS testbed and clinical datasets, our approach achieves a 92.8% hazard mitigation success rate—improving over pure rule-based or pure data-driven baselines by >76%. It guarantees zero false negatives, maintains low false positive rates, and introduces no new safety risks.
This study addresses the challenge of providing certifiable runtime safety guarantees prior to tool invocation, focusing on three core issues: the representability of policy states, the observability of monitoring evidence, and the impact of interventions on future behavior. To this end, we propose the first formal theoretical framework for runtime safety-executable boundaries, distinguishing among static policy executability, statistical calibration under exogenous legal constraints, and closed-loop intervention effects. Building upon finitely controlled models, we develop a method for closed-loop safety certification that integrates register model identification, Neyman–Pearson hypothesis testing, conformal calibration, and occupancy planning. Empirical validation through static diagnosis, model enumeration, representation rewriting, and closed-loop re-execution experiments demonstrates the efficacy of our approach and exposes the fundamental limitations of static calibration under representation attacks.
Traditional static safety cases struggle to dynamically respond to runtime evidence and cannot continuously quantify confidence in system safety. This work proposes a dynamic safety argumentation framework grounded in subjective logic, which integrates design-time evidence with runtime Safety Performance Indicators (SPIs). By employing a sliding window mechanism to process SPI data in real time, the framework introduces a confidence-updating rule prioritizing safety responsiveness—gradually increasing confidence in the absence of violations while imposing swift penalties upon detection of anomalies—thereby overcoming limitations inherent in conventional Bayesian posterior updating. The approach is validated through simulations of an assistive function in construction zones, effectively demonstrating the dynamic evolution of confidence in a machine learning–driven traffic cone detection component as informed by runtime evidence.
This study addresses the dual challenges of functional safety and Safety of the Intended Functionality (SOTIF) in high-level automated driving systems operating in dynamic, unknown environments, where reliable performance is difficult to ensure under system failures or exposure to scenarios absent from training data. The authors propose a hierarchical fault-tolerant architecture that, for the first time, integrates a functional monitor—based on voting consensus among multi-channel heterogeneous AI perception modules—and an anomaly monitor designed to detect unknown or novel objects within a unified framework. This integration enables runtime-triggered minimal-risk maneuvers, safe degradation, and data logging, thereby establishing a closed-loop safety enhancement mechanism spanning development and operational phases. Real-world vehicle tests demonstrate that this dual-monitor approach significantly improves system robustness and safety in unseen scenarios while complying with ISO 26262 and SOTIF requirements.
This work proposes CopilotVerifier, an automated verification framework designed to enhance the correctness and trustworthiness of runtime monitoring code in safety-critical systems by complementing the Copilot compiler. CopilotVerifier is the first to decompose the bisimulation relation between source programs and their compiled C code into verifiable conditions. By integrating symbolic execution (via Crucible) with SMT solving (through What4), the framework automatically generates formal proofs that guarantee semantic equivalence—ensuring identical outputs and consistent crash behaviors under equivalent inputs. This approach significantly strengthens compiler assurance with modest computational overhead and lays the groundwork for producing human-auditable formal arguments of correctness.