π€ AI Summary
This work addresses the challenge of optimizing objective functions whose gradients are not analytically tractableβa common scenario in training and fine-tuning generative models and learning with latent variables. The authors propose a class of stochastic particle optimization methods grounded in diffusion processes, leveraging mean-field dynamics and their interacting particle approximations to construct a unified framework for optimizing parameterized distributions via integral-form gradients. This framework encompasses and generalizes several existing algorithms. Theoretically, the study establishes non-asymptotic error bounds for the continuous-time particle system and proves exponential convergence under a joint contractivity assumption. Algorithmically, it integrates momentum, higher-order Langevin variants, and numerical schemes for stochastic differential equations. Empirical results demonstrate significant improvements over baseline methods in maximum marginal likelihood estimation and energy-based model training.
π Abstract
We develop a class of diffusion-based stochastic particle optimisation methods for loss functions with intractable gradients. Specifically, we consider problems in which the loss gradient is an integral with respect to a parameter-dependent distribution, a structure that includes training generative models, fine-tuning, and learning latent-variable models. We introduce mean-field dynamics and its interacting-particle approximations, which contain several existing algorithms as special cases and provides a route to constructing new methods. Under well-posedness and joint contractivity assumptions, we prove exponential convergence and show that the continuous-time particle system admits a non-asymptotic error bound. We illustrate it by developing momentum and higher-order Langevin variants and evaluating them on maximum marginal-likelihood estimation and energy-based-model training.