๐ค AI Summary
This paper unifies three dominant paradigms in optimal control and motion planningโModel Predictive Path Integral (MPPI) control, reinforcement learning (RL), and diffusion models. Methodologically, it leverages gradient optimization over the Gibbs measure to establish rigorous theoretical connections. The contributions are threefold: (i) MPPI is proven equivalent to gradient ascent on the Gibbs energy functional; (ii) under fixed initial states, policy gradient methods reduce exactly to MPPI; and (iii) the reverse sampling update rule of diffusion models coincides identically with the MPPI update. Collectively, these results establish a fundamental mathematical equivalence among the three frameworks. Beyond unification, the analysis reveals shared mechanistic principles underlying generative planning methods and enables the design of robust, efficient, and interpretable unified generative optimal controllers. This work provides a novel paradigm for tightly integrating learning-based and model-based control.
๐ Abstract
Model Predictive Path Integral (MPPI) control, Reinforcement Learning (RL), and Diffusion Models have each demonstrated strong performance in trajectory optimization, decision-making, and motion planning. However, these approaches have traditionally been treated as distinct methodologies with separate optimization frameworks. In this work, we establish a unified perspective that connects MPPI, RL, and Diffusion Models through gradient-based optimization on the Gibbs measure. We first show that MPPI can be interpreted as performing gradient ascent on a smoothed energy function. We then demonstrate that Policy Gradient methods reduce to MPPI when treating policy parameters as control variables under a fixed initial state. Additionally, we establish that the reverse sampling process in diffusion models follows the same update rule as MPPI.