🤖 AI Summary
This study addresses the instability and inefficiency of initial noise optimization in black-box reward scenarios. To overcome these challenges, this work proposes ZeNOVA, a framework that achieves gradient-free alignment through manifold-constrained hyperspherical Langevin dynamics. By incorporating annealed soft value guidance and a Metropolis-Hastings jump mechanism, while leveraging the geometric properties of Gaussian priors, the proposed method effectively overcomes the stability bottlenecks inherent in zeroth-order optimization. Extensive experiments on image and video generation tasks demonstrate that ZeNOVA significantly outperforms existing baselines, achieving more stable and efficient high-reward optimization of initial noise.
📝 Abstract
Recent advances in distillation and flow-map models have enabled deterministic one- or few-step generation for high-quality data, facilitating a new branch of reward alignment approaches that directly optimize the initial noise from a Gaussian distribution. However, most existing initial-noise optimization methods rely on first-order gradient information, which is either inapplicable or suffers from instability and inefficiency in black-box reward scenarios. Here, we introduce ZeNOVA, a stable and efficient initial noise alignment method in a gradient-free manner. Specifically, we address existing algorithms' major challenge in black-box scenarios through annealed soft-value guidance, manifold-constrained hyperspherical Langevin dynamics, and Metropolis-Hastings jumping. Extensive experiments on image and video generative models show that ZeNOVA outperforms all evaluated zeroth-order baselines by optimizing the initial noise toward higher rewards substantially more stably while exploiting the geometry of the Gaussian prior, demonstrating its practical applicability to various black-box reward alignment.