Diffusion Policy Improvement with Proposal-Conditioned Refinement Flows

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of conventional diffusion policies in offline reinforcement learning, where KL penalties suppress high-value yet low-density actions and existing methods struggle to model multimodal behaviors. To overcome these challenges, this work proposes PReFlow, a framework integrating critic-based proposal selection with conditional refinement flows. The method introduces a proposal-centric Gaussian reference distribution to regularize action variations and employs a simulation-free, closed-form adjoint matching objective to circumvent backward adjoint computation, while optimizing the Gibbs policy via KL regularization. Evaluated across 50 OGBench tasks, PReFlow demonstrates superior performance, achieving the highest aggregate score following online fine-tuning and reaching a 91% success rate within 500,000 interaction steps.
📝 Abstract
Diffusion and flow policies can model complex behaviors in offline reinforcement learning (RL). However, penalizing their KL divergence from the behavior policy can discourage actions having high critic values with low behavior density. Directly refining behavior proposals may be an alternative, yet Gaussian or deterministic editors limit expressiveness to represent multiple separated modes for the same proposal. In this work, we introduce Proposal-Conditioned Refinement Flows (PReFlow), a policy extraction method combining critic-based proposal selection with a conditional refinement flow. To optimize proposal selection and refinement together, we formulate a KL-regularized objective whose optimum induces a Gibbs policy over final actions under a Gaussian-smoothed behavior prior. The refinement flow can represent multiple high value modes, while a proposal-centered Gaussian reference regulates large action changes. This Gaussian reference further enables us to make use of simulation-free, closed form adjoint matching targets from sampled endpoints and critic gradients, yielding a single velocity regression loss without a backward adjoint solve. On 50 OGBench tasks, PReFlow achieves competitive offline performance and the highest aggregate score among the compared methods after online fine-tuning, reaching 91\% after 500K environment steps.
Problem

Research questions and friction points this paper is trying to address.

Offline Reinforcement Learning
Diffusion Policy
KL Divergence Penalty
Policy Refinement
Multi-modal Representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Proposal-Conditioned Refinement Flows
Offline Reinforcement Learning
KL-Regularized Objective
Adjoint Matching
Flow Policy
🔎 Similar Papers
2024-07-16arXiv.orgCitations: 2
💼 Related Jobs
No related jobs found.