Value-at-Risk Constrained Policy Optimization

📅 2026-01-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the problem of safe policy optimization under Value-at-Risk (VaR) constraints in reinforcement learning. To prevent constraint violations during training, the authors propose a sample-efficient conservative policy optimization method that directly incorporates VaR constraints into the policy optimization framework. By leveraging a one-sided Chebyshev inequality, they construct a differentiable surrogate constraint based on the first and second moments of the cost return, which is then integrated with an extended trust region mechanism. This approach provides theoretical guarantees for both policy improvement and worst-case constraint satisfaction. Empirical results demonstrate that the method achieves zero constraint violations throughout training across multiple environments, significantly outperforming existing baselines while maintaining high sample efficiency.

Technology Category

Constraint Satisfaction and Optimization: Constraint OptimizationReasoning under Uncertainty: Stochastic OptimizationIntelligent Robots: Learning & Optimization for ROB

Application Category

Responsible Web: Human-perceived consequences of algorithmic deployment on the webSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingEconomics, Online Markets and Human Computation: Trust and reliance of crowd workers and data experts on GenAI
📝 Abstract
We introduce the Value-at-Risk Constrained Policy Optimization algorithm (VaR-CPO), a sample efficient and conservative method designed to optimize Value-at-Risk (VaR) constraints directly. Empirically, we demonstrate that VaR-CPO is capable of safe exploration, achieving zero constraint violations during training in feasible environments, a critical property that baseline methods fail to uphold. To overcome the inherent non-differentiability of the VaR constraint, we employ the one-sided Chebyshev inequality to obtain a tractable surrogate based on the first two moments of the cost return. Additionally, by extending the trust-region framework of the Constrained Policy Optimization (CPO) method, we provide rigorous worst-case bounds for both policy improvement and constraint violation during the training process.
Problem

Research questions and friction points this paper is trying to address.

Value-at-Risk
constrained policy optimization
safe exploration
risk constraint
reinforcement learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Value-at-Risk
Constrained Policy Optimization
Safe Reinforcement Learning
Chebyshev Inequality
Trust-Region Method
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Rohan Tangri
Machine Learning Research Group, University of Oxford, Oxford, United Kingdom
Jan-Peter Calliess
Jan-Peter Calliess
Oxford University
Machine LearningControlFinanceArtificial IntelligenceNumerical Mathematics