🤖 AI Summary
This work addresses the problem of safe policy optimization under Value-at-Risk (VaR) constraints in reinforcement learning. To prevent constraint violations during training, the authors propose a sample-efficient conservative policy optimization method that directly incorporates VaR constraints into the policy optimization framework. By leveraging a one-sided Chebyshev inequality, they construct a differentiable surrogate constraint based on the first and second moments of the cost return, which is then integrated with an extended trust region mechanism. This approach provides theoretical guarantees for both policy improvement and worst-case constraint satisfaction. Empirical results demonstrate that the method achieves zero constraint violations throughout training across multiple environments, significantly outperforming existing baselines while maintaining high sample efficiency.
📝 Abstract
We introduce the Value-at-Risk Constrained Policy Optimization algorithm (VaR-CPO), a sample efficient and conservative method designed to optimize Value-at-Risk (VaR) constraints directly. Empirically, we demonstrate that VaR-CPO is capable of safe exploration, achieving zero constraint violations during training in feasible environments, a critical property that baseline methods fail to uphold. To overcome the inherent non-differentiability of the VaR constraint, we employ the one-sided Chebyshev inequality to obtain a tractable surrogate based on the first two moments of the cost return. Additionally, by extending the trust-region framework of the Constrained Policy Optimization (CPO) method, we provide rigorous worst-case bounds for both policy improvement and constraint violation during the training process.