CORE-RL: Confidence-Oriented Reliability Evaluation of Black-Box Reinforcement Learning Policies

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of evaluating multi-objective alignment and environmental robustness in black-box reinforcement learning policies whose internal mechanisms remain inaccessible. We propose CORE-RL, an evaluation framework that introduces a unified reliability metric integrating task termination and safety violations to prevent failure masking. The framework constructs certified envelopes through perception/actuation noise injection and environmental perturbation testing, while employing Clopper-Pearson and Hoeffding bounds to ensure rigorous statistical inference under finite-sample conditions. Experimental results demonstrate that CORE-RL can automatically reject non-compliant policies and precisely delineate safe operational design domains. By providing a quantitative, transparent, and reproducible evaluation basis, this work offers a principled approach for the safe deployment of black-box reinforcement learning policies in complex environments.
📝 Abstract
The deployment of Reinforcement Learning (RL) agents in critical domains must be preceded with a pipeline to evaluate the alignment of the RL agent with complex multi-objective specifications and robustness under real-world environmental drift. However, to protect intellectual property, the RL agent may be delivered for evaluation as opaque executable or remote API, which makes traditional evaluation techniques based on the internals of the policies infeasible. To address this gap, CORE-RL: Confidence-Oriented Reliability Evaluation of black box RL policy is proposed in this paper. The CORE-RL pipeline introduces a Unified Reliability Metric that formally integrates early task termination and safety constraint violations, preventing unsafe policies from masking failures through premature episode halts. By subjecting the policy to a noise certification envelope of perceptual noise, actuation noise and change in environment dynamics, the pipeline computes the finite-sample Clopper-Pearson bounds on unified reliability metric and Hoeffdings'lower bound on reward and safety cost. The pipeline then defines safe operational design domain to report high-confidence certificates for safety and expected performance. Experiments on continuous control tasks demonstrate the CORE-RL pipeline's ability to automatically reject non-compliant policies and map the safe Operational Design Domain (ODD) of safety-aware policies. Thus CORE-RL provides an evaluation framework towards a quantitative, transparent and reproducible, statistical rationale necessary to safely evaluate, compare, and deploy black box RL solutions.
Problem

Research questions and friction points this paper is trying to address.

Black-box Reinforcement Learning
Reliability Evaluation
Safety Certification
Operational Design Domain
Robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Black-box Reinforcement Learning
Reliability Evaluation
Unified Reliability Metric
Operational Design Domain
Statistical Bounds
S
Santhosh GS
Center for Responsible AI, IIT Madras, Chennai, India
A
Ananya Ravi
Center for Responsible AI, IIT Madras, Chennai, India
D
Devika Jay
Center for Responsible AI, IIT Madras, Chennai, India
S
Saurav Prakash
Center for Responsible AI, IIT Madras, Chennai, India
Balaraman Ravindran
Balaraman Ravindran
Professor of Data Science and AI, Wadhwani School of Data Science and AI, IIT Madras
Reinforcement LearningData MiningNetwork AnalysisResponsible AI
A
Abhishek Sarkar
Ericsson Research, India
P
Perepu Satheesh Kumar
Ericsson Research, India
K
Kaushik Dey
Ericsson Research, India