🤖 AI Summary
This work addresses the challenge of simultaneously achieving high performance and formal guarantees of safety and robustness in learning-based control for safety-critical cyber-physical systems. The paper proposes SMC-ES, a novel framework that deeply integrates evolutionary strategies with statistical model checking to enable provably sound, high-confidence verification of control policies with respect to performance, safety, and robustness. Evaluated on Gymnasium and Safety-Gymnasium benchmarks, SMC-ES attains performance comparable to state-of-the-art model-free deep reinforcement learning (DRL) and safe DRL methods, while incurring only moderate computational overhead. Crucially, it provides rigorous formal guarantees on violation probabilities, thereby bridging a critical gap in the field by offering verifiable assurance for learned control policies that has been largely absent in existing approaches.
📝 Abstract
The deployment of autonomous cyber-physical systems in safety-critical environments requires closed-loop control strategies (i.e., policies) that are not only performant but also provably safe and robust. While learning-based methodologies such as Reinforcement Learning offer flexible and scalable approaches to automatically synthesize such controllers, they typically lack the formal guarantees necessary for safe deployment. To bridge this gap, we propose a novel simulation-based methodology to automatically synthesize policies with formal guarantees regarding performance, safety, and robustness specifications. Specifically, given a set of properties to verify, a confidence parameter $δ$ and an allowable failure probability $\varepsilon$, our method guarantees that the synthesized policy comes with a certificate: with confidence at least $1 - δ$, the probability of encountering a scenario where the given properties are violated is at most $\varepsilon$. We demonstrate the feasibility of our approach by developing SMC-ES, an algorithm that integrates Evolutionary Strategies with Statistical Model Checking-based verification. We evaluate SMC-ES on a suite of continuous control tasks using Gymnasium and Safety Gymnasium testbeds. Results show that, at the price of a sustainable increase in computational cost, our algorithm provides formal guarantees regarding performance, safety, and robustness specifications, while performing competitively against leading model-free Deep Reinforcement Learning (DRL) and Safe-DRL baselines.