Convex-Concave Reinforcement Learning

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the non-convexity of the return maximization objective in reinforcement learning and the reliance of existing methods on surrogate approximations. By operating in log-density ratio coordinates, this work reveals that the original problem is fundamentally a difference-of-convex (DC) program subject to DC constraints. It exactly solves the primal objective through decision-wise importance sampling and sequential convex programming, thereby unifying classical algorithms such as Natural Policy Gradient (NPG) while introducing a multi-step coupling mechanism. Experimental results demonstrate that the proposed approach significantly outperforms Proximal Policy Optimization (PPO) in long-horizon credit assignment and healthcare tasks, achieving faster convergence and an 11.3% improvement in the area under the training curve.
📝 Abstract
Policy learning drives many of the most consequential and heavily-invested applications of reinforcement learning today. Yet the core optimization problem it rests on (maximizing expected return) is notoriously non-convex, even under a direct policy parameterization, and the field has largely responded by avoiding it: optimizing convex surrogate approximations of the return under trust-region constraints (NPG, TRPO, PPO, AWR). We show that this seemingly unstructured problem is not actually structureless. In log-density-ratio coordinates $y := \log[π/π_n]$, the exact per-iteration objective, computable via per-decision importance sampling (PDIS), is a difference-of-convex-constrained difference-of-convex (DC-constrained DC) program. This structure lets us move beyond surrogate approximations: it recovers CPI, NPG, TRPO, and AWR as special cases along interpretable axes, and it opens a multi-step axis $k$ that couples consecutive decisions. We solve the per-iteration program with sequential convex programming (SCP), the standard solver for difference-of-convex problems, and give convergence guarantees under mild conditions, bridging the difference-of-convex optimization and RL literatures. Empirically, multi-step Convex-Concave RL (CCRL) wins on diagnostic MDPs where credit must propagate across a horizon (its advantage growing with the dependency length), is competitive with a tuned PPO on classic control, and on a realistic, stochastic, mid-horizon healthcare domain converges markedly faster than tuned PPO to the same near-optimal survival, with an 11.3% higher area under the training curve.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Policy Optimization
Non-convex Optimization
Surrogate Approximation
Difference-of-Convex
Innovation

Methods, ideas, or system contributions that make the work stand out.

Convex-Concave Reinforcement Learning
Difference-of-Convex Programming
Sequential Convex Programming
Policy Optimization
Per-Decision Importance Sampling