Reachability-Aware Diffusion Policy Optimization

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation that existing safe policy optimization methods in continuous control predominantly rely on explicit dynamics models, restricting their applicability in model-free scenarios. To overcome this, we propose RADPO, a framework for safe diffusion policy optimization that eliminates the need for learned dynamics models, action gradients, or backpropagation through differentiation. Specifically, RADPO integrates predictive First Hitting Time (FHIT) reachability value estimation with cumulative cost budget feedback. It reshapes rewards via discounted reachability values and introduces a dynamic dual multiplier adjustment mechanism to synergistically optimize the diffusion policy. Evaluated across ten benchmark tasks, RADPO achieves competitive reward-cost trade-offs while significantly reducing constraint violation rates, thereby establishing an efficient new paradigm for model-free safe reinforcement learning.
📝 Abstract
Diffusion policies provide expressive action distributions for continuous-control reinforcement learning. However, safety-aware online diffusion policy optimization remains underexplored, particularly methods that use predictive reachability information without an explicit dynamics model. We propose Reachability-Aware Diffusion Policy Optimization (RADPO), a model-free method that combines predictive first-hit safety estimation with cumulative-cost budget feedback. RADPO learns a discounted first-hit reachability value that captures the discounted risk of a cost event, assigns larger weight to events that occur sooner, and uses this signal to shape the reward. A separate dual-like multiplier adjusts the shaping strength according to realized episodic costs relative to a prescribed budget. The diffusion actor improves through weighted denoising regression on candidate actions scored by the reward critic. Our approach requires neither a learned dynamics model, action gradients through the critics, nor differentiation through the reverse diffusion sampler. We establish theoretical properties of the reachability value and show that accumulated reachability penalty provides a conservative surrogate for future discounted cumulative cost. Across ten continuous-control safety tasks, RADPO achieves competitive reward-cost trade-offs, with substantial reductions in constraint violations on several tasks relative to the compared baselines. Our theoretical and empirical analysis supports that combining reachability with cumulative budget feedback is a viable approach to safety-aware diffusion policies.
Problem

Research questions and friction points this paper is trying to address.

Diffusion Policy
Safe Reinforcement Learning
Reachability
Continuous Control
Model-free
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diffusion Policy
Reachability-Aware
Model-Free
Safety-Constrained Reinforcement Learning
Reward Shaping
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hikmet Simsir
Department of Computer Engineering, Bilkent University
K
Kutay Demiray
Department of Computer Engineering, Bilkent University
Ozgur S. Oguz
Ozgur S. Oguz
Asst. Prof. @Bilkent University
Artificial IntelligenceRoboticsTask and Motion PlanningHRI