🤖 AI Summary
Diffusion policies in continuous control suffer from high computational overhead due to iterative denoising. To address this, this work proposes the Prefix-Optimal Generation Policy (POGP), which introduces Bellman-style recursion into the diffusion denoising chain for the first time, yielding a prefix value function. This function serves both as an auxiliary training objective to enhance the quality of intermediate actions and enables adaptive early stopping during inference. Evaluated on the MuJoCo benchmark suite, POGP reduces the number of denoising steps by an average factor of 2.7 while maintaining near-optimal performance and achieves approximately a 3.5% improvement in final return over existing dynamic diffusion methods.
📝 Abstract
Diffusion policies are a powerful policy class for continuous control, but their iterative denoising process creates a substantial computational bottleneck. Reducing this cost requires adapting the number of denoising steps to the difficulty of each action while preserving task performance. We introduce Prefix-Optimal Generative Policies (POGP), a framework that learns a prefix value function at every intermediate denoising step through a Bellman-style recursion over the denoising chain. The prefix value function serves two purposes: it provides an auxiliary training objective that encourages intermediate outputs to become high-quality actions, and it enables a test-time stopping rule that terminates denoising when additional steps are unlikely to produce meaningful improvement. Across four MuJoCo environments and comparisons with 12 baselines, POGP reduces the required number of denoising iterations by approximately 2.7-fold while retaining near-full task performance. Compared with state-of-the-art dynamic diffusion baselines, prefix training also improves final task performance by approximately 3.5%. These results indicate that supervising intermediate denoising steps is useful not only for adaptive early stopping, but also as an auxiliary objective that improves the learned policy.