Schr\"odinger--F\"ollmer Actor--Critic: Diffusion Policy Improvement with Finite-Sample Analysis

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of advantage-weighted updates in diffusion policies, which cannot explicitly specify the sampling mechanism for target distributions. To overcome this, we propose SFAC, a method that formulates policy updates as drift corrections via the Doob h-transform and leverages conditional diffusion to achieve KL-regularized offline-to-online policy improvement. Furthermore, SFAC integrates a minimax Bellman critic, self-normalized importance sampling, and supervised regression techniques, accompanied by derived finite-sample error bounds. Experimental results demonstrate that SFAC significantly improves final returns on continuous control tasks, outperforming existing offline initialization baselines.
📝 Abstract
Diffusion policies represent multimodal action distributions, but an advantage-weighted update does not specify how to sample from the resulting target distribution. We propose Schr\"odinger--F\"ollmer Actor--Critic (SFAC), an offline-to-online reinforcement learning (RL) method for Kullback--Leibler (KL)-regularized policy improvement through conditional diffusion. A minimax Bellman critic estimates the advantage function, which defines an exponentially tilted target policy. A Doob $h$-transform expresses this update as a correction to the reference diffusion drift. We derive a posterior-mean representation of the correction and estimate it using paired self-normalized importance sampling (SNIS). Supervised regression on these drift targets updates the neural actor without critic action gradients. In the small-update regime, the KL-regularized update follows the natural policy-gradient direction, and the Doob correction represents the same local change in the space of diffusion drifts. Under suitable conditions, we derive finite-sample bounds that separate the effects of critic estimation, neural drift regression, finite-sample SNIS, diffusion discretization, and inherited actor error on expected average policy suboptimality. Synthetic experiments assess the accuracy of approximation to prescribed advantage-tilted targets and sensitivity to sampling budgets. On six offline-to-online continuous-control tasks, a reference-anchored implementation achieves higher final-window returns than those of its corresponding offline initialization.
Problem

Research questions and friction points this paper is trying to address.

diffusion policies
policy improvement
advantage-weighted update
target distribution sampling
offline-to-online reinforcement learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diffusion Policy
Actor-Critic
Doob h-transform
Finite-Sample Analysis
Offline-to-Online Reinforcement Learning
🔎 Similar Papers
No similar papers found.
Yuling Jiao
Yuling Jiao
University of Wuhan
Deep learningScientific and statistical computingInverse problem
L
Lican Kang
Institute for Math and AI, Hubei Key Laboratory of Computational Science, and School of Artificial Intelligence, Wuhan University, Wuhan, 430072, China
J
Jerry Zhijian Yang
School of Mathematics and Statistics, and Hubei Key Laboratory of Computational Science, Wuhan University, Wuhan, 430072, China
J
Jincheng Ying
School of Mathematics and Statistics, Wuhan University, Wuhan, 430072, China