PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the difficulty of diffusion models in precisely generating compositional prompts, such as those involving quantities and spatial relationships. To this end, it proposes a training-free test-time search method that reformulates diffusion sampling as intermediate latent search guided by a multimodal critic. Departing from conventional paradigms that merely select from final samples, this approach leverages natural language feedback to edit and branch denoising trajectories in real time. The core technique integrates multimodal critic scoring, semantic prompt editing, and local re-noised latent continuation. Experimental results demonstrate that the proposed method significantly outperforms budget-matched Best-of-N and scalar search baselines across both image and video benchmarks.
📝 Abstract
Diffusion models can produce striking images and videos, but they still struggle with the compositional details that make a generation faithful to a prompt, such as object counts, attribute binding, spatial relations, and temporally grounded actions. A common way to improve prompt satisfaction is to spend more compute at test time through Best-of-N sampling, but final-sample selection is fixed. Best-of-N can only choose among completed outputs and cannot repair a promising trajectory before it fails. We introduce PreviewDiff, a training-free test-time search method that turns diffusion sampling from scalar search into a multimodal critic-guided search over intermediate latents. At selected denoising checkpoints, PreviewDiff decodes a partial preview, asks a multimodal judge to score and critique it, and uses the resulting natural-language feedback to branch over semantic prompt edits and locally re-noised latent continuations. These branches are then scored and selectively rolled forward, allowing verifier compute to guide generation while the sample is still editable. Across image and video generation benchmarks, PreviewDiff consistently improves over budget-matched Best-of-N selection and strong scalar-search baselines. Ablations show that earlier interventions and increased search width provide the largest gains, while deeper search and additional semantic variants offer complementary improvements. PreviewDiff demonstrates that multimodal feedback is most useful not only as a final verifier, but as an active controller inside the denoising process.
Problem

Research questions and friction points this paper is trying to address.

Diffusion models
Compositional generation
Test-time search
Prompt faithfulness
Best-of-N sampling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-time search
Multimodal critic-guided generation
Diffusion latents
Training-free
Intermediate preview feedback
🔎 Similar Papers
No similar papers found.