LatticeSMC: Where to Spend Inference-Time Compute in Chunked Sequence Generators

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of uneven inference budget allocation and inconsistent guidance criteria in chunked sequence generation by proposing a budget-matched chunked guidance framework. Based on a two-dimensional lattice Feynman-Kac model, we design a sampler that optimizes resampling strategies across denoising steps and chunk indices. We derive a scaling theorem to establish design precision, enabling prefix-evaluable reward twisting without additional estimators, and formalize the theory of optimal resampling. Experimental results demonstrate that music-dance alignment improves from 0.234 to 0.441, with prompt adherence reaching 0.560. These outcomes significantly surpass Best-of-N baselines and are further corroborated by human preference evaluations.
📝 Abstract
Long-form generators for music, motion and video produce sequences chunk by chunk, with each chunk generated by iterative denoising while rewards are defined over the full sequence. Existing inference-time steering methods typically act on one axis at a time: best-of-N at the end, Feynman-Kac steering across denoising steps, or streaming pruning across chunks, and are often compared under unmatched compute or different return rules. We introduce budget-matched chunked steering and propose LatticeSMC, a sampler derived from a Feynman-Kac model on the two-dimensional lattice of chunk index and denoising step. Two telescoping results make its design exact: for chunk-additive rewards, the two axes induce identical weights, so resampling should occur where lookahead is cheapest; for terminal rewards, any prefix score defines an exact intermediate potential, making prefix-evaluable rewards twists with no estimation or extra denoiser calls. LatticeSMC resamples on these potentials at chunk boundaries and, when scoring is free, within chunks, returning either a weighted draw or the best particle. Under matched compute, on music-to-dance diffusion and 40-second text-to-music generation, it raises beat alignment from 0.234 to 0.441 (best-of-N: 0.354) and prompt adherence from 0.470 to 0.560 at 32 particles, while preserving held-out quality. It also retains its advantage on long-range rewards and is preferred by human raters in 60-77 percent of pairwise comparisons. Finally, we show that commitment strength should follow the information in the current potential, while the value of lookahead is predicted by the within-set predictability of future reward.
Problem

Research questions and friction points this paper is trying to address.

chunked sequence generation
inference-time compute allocation
long-form generation
iterative denoising
sequential Monte Carlo
Innovation

Methods, ideas, or system contributions that make the work stand out.

LatticeSMC
Feynman-Kac model
chunked sequence generation
inference-time compute allocation
budget-matched steering
🔎 Similar Papers
No similar papers found.