Best-of-$N$ Guidance for Test-time Diffusion Alignment

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of traditional Best-of-N (BoN) sampling in test-time alignment for diffusion models, which fails to leverage reward signals for optimizing generation trajectories and yields suboptimal average quality. To overcome this, we propose BoNG, a method that integrates the BoN principle directly into the reverse diffusion process. BoNG introduces an innovative mechanism featuring online selection and asymmetric guided interaction among denoising particles, transcending end-stage filtering constraints to drive sampling trajectories toward high-reward regions. Experimental results demonstrate that BoNG achieves state-of-the-art performance in 29 out of 36 comparisons (80.56%), attains ImageReward scores 1.3 times higher than existing methods, and simultaneously delivers a 1.6-fold acceleration.
📝 Abstract
Diffusion models achieve strong generative performance but often struggle to align generated samples with human preferences measured by a reward model. A simple yet effective algorithm for test-time alignment is Best-of-$N$ (BoN) sampling, which draws $N$ i.i.d. samples from a pre-trained diffusion model and outputs the single highest-reward sample. Despite its empirical success, BoN makes limited use of reward information, as it is incorporated only at the final selection stage without influencing the reverse diffusion trajectory during sampling. Consequently, BoN sampling does not improve the average alignment of generated samples and is primarily suited to single-output settings. We propose Best-of-$N$ Guidance (BoNG), a novel method that integrates the principle of BoN sampling directly into the reverse diffusion process. BoNG performs online BoN selection over denoising particles and adjusts the reverse diffusion process to steer the particle population toward higher-reward regions during generation. Specifically, by introducing an asymmetric guidance interaction among denoising particles, BoNG uses the current BoN particle as a guidance signal to the rest of the particle population. This particle-level interaction reshapes the sampling process toward higher-reward regions, enabling BoNG to improve not only the final best sample beyond Vanilla BoN sampling, but also the average quality of generated samples. Over 36 empirical comparisons, BoNG achieves the best performance in 29 cases, ranking first in 80.56% of the comparisons against SMC and Vanilla BoN sampling. BoNG also supports multi-output capability, achieving 1.3$\times$ ImageReward score of the latest sample-based guidance method with a 1.6$\times$ speedup. We release the code at https://github.com/aailab-kaist/BoNG.
Problem

Research questions and friction points this paper is trying to address.

Diffusion models
Test-time alignment
Best-of-N sampling
Human preferences
Reward model
Innovation

Methods, ideas, or system contributions that make the work stand out.

Diffusion Models
Test-time Alignment
Best-of-N Guidance
Particle Interaction
Reward Model
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Richard Lee Kim
KAIST
Yeongmin Kim
Yeongmin Kim
KAIST
Generative ModelsMachine Learning
G
Gyuwon Sim
KAIST
T
Taekyu Kim
KAIST
M
Minsang Park
KAIST
I
Il-chul Moon
KAIST, summary.ai