🤖 AI Summary
This study addresses the limitation that existing safe policy optimization methods in continuous control predominantly rely on explicit dynamics models, restricting their applicability in model-free scenarios. To overcome this, we propose RADPO, a framework for safe diffusion policy optimization that eliminates the need for learned dynamics models, action gradients, or backpropagation through differentiation. Specifically, RADPO integrates predictive First Hitting Time (FHIT) reachability value estimation with cumulative cost budget feedback. It reshapes rewards via discounted reachability values and introduces a dynamic dual multiplier adjustment mechanism to synergistically optimize the diffusion policy. Evaluated across ten benchmark tasks, RADPO achieves competitive reward-cost trade-offs while significantly reducing constraint violation rates, thereby establishing an efficient new paradigm for model-free safe reinforcement learning.
📝 Abstract
Diffusion policies provide expressive action distributions for continuous-control reinforcement learning. However, safety-aware online diffusion policy optimization remains underexplored, particularly methods that use predictive reachability information without an explicit dynamics model. We propose Reachability-Aware Diffusion Policy Optimization (RADPO), a model-free method that combines predictive first-hit safety estimation with cumulative-cost budget feedback. RADPO learns a discounted first-hit reachability value that captures the discounted risk of a cost event, assigns larger weight to events that occur sooner, and uses this signal to shape the reward. A separate dual-like multiplier adjusts the shaping strength according to realized episodic costs relative to a prescribed budget. The diffusion actor improves through weighted denoising regression on candidate actions scored by the reward critic. Our approach requires neither a learned dynamics model, action gradients through the critics, nor differentiation through the reverse diffusion sampler. We establish theoretical properties of the reachability value and show that accumulated reachability penalty provides a conservative surrogate for future discounted cumulative cost. Across ten continuous-control safety tasks, RADPO achieves competitive reward-cost trade-offs, with substantial reductions in constraint violations on several tasks relative to the compared baselines. Our theoretical and empirical analysis supports that combining reachability with cumulative budget feedback is a viable approach to safety-aware diffusion policies.