🤖 AI Summary
This study addresses the limitation of advantage-weighted updates in diffusion policies, which cannot explicitly specify the sampling mechanism for target distributions. To overcome this, we propose SFAC, a method that formulates policy updates as drift corrections via the Doob h-transform and leverages conditional diffusion to achieve KL-regularized offline-to-online policy improvement. Furthermore, SFAC integrates a minimax Bellman critic, self-normalized importance sampling, and supervised regression techniques, accompanied by derived finite-sample error bounds. Experimental results demonstrate that SFAC significantly improves final returns on continuous control tasks, outperforming existing offline initialization baselines.
📝 Abstract
Diffusion policies represent multimodal action distributions, but an advantage-weighted update does not specify how to sample from the resulting target distribution. We propose Schr\"odinger--F\"ollmer Actor--Critic (SFAC), an offline-to-online reinforcement learning (RL) method for Kullback--Leibler (KL)-regularized policy improvement through conditional diffusion. A minimax Bellman critic estimates the advantage function, which defines an exponentially tilted target policy. A Doob $h$-transform expresses this update as a correction to the reference diffusion drift. We derive a posterior-mean representation of the correction and estimate it using paired self-normalized importance sampling (SNIS). Supervised regression on these drift targets updates the neural actor without critic action gradients. In the small-update regime, the KL-regularized update follows the natural policy-gradient direction, and the Doob correction represents the same local change in the space of diffusion drifts. Under suitable conditions, we derive finite-sample bounds that separate the effects of critic estimation, neural drift regression, finite-sample SNIS, diffusion discretization, and inherited actor error on expected average policy suboptimality. Synthetic experiments assess the accuracy of approximation to prescribed advantage-tilted targets and sensitivity to sampling budgets. On six offline-to-online continuous-control tasks, a reference-anchored implementation achieves higher final-window returns than those of its corresponding offline initialization.