🤖 AI Summary
This study addresses the limitation of unidirectional inference in discrete diffusion models, which prevents the correction of erroneous tokens during protein generation. To overcome this, we propose the Spectral Feedback algorithm, endowing protein generative models with iterative self-correction capabilities. The method shifts the alignment focus from token labels to edit position selection, employing a feedback loop to localize and mask-resample target positions for editing. Furthermore, it leverages the sparse Fourier transform to efficiently optimize a complex, interdependent value function over edit sets. Evaluated on the protein inverse folding task, our approach increases the proportion of stable proteins generated by pretrained models by 32.3%, significantly outperforming existing methods without requiring modifications to the underlying generative process.
📝 Abstract
Reward maximization alignment methods for discrete diffusion models have primarily focused on steering the reverse process, either by influencing token logits or by selecting favorable sequences at intermediate steps. These approaches largely treat inference as a unidirectional process, lacking mechanisms for revisiting undesirable token selections. We introduce Spectral Feedback, an algorithm that selects edit-positions in a feedback loop, allowing the model to iteratively correct its own generations. This approach leverages the mask structure of discrete diffusion models by re-masking and re-sampling tokens, analogous to image editing methods that reintroduce noisy latents and re-run the reverse process. While prior alignment methods focus on what token labels to assign to maximize a target reward, we instead treat which tokens to revisit as the central alignment problem. Selecting edit-positions is challenging because edit effects are interdependent: the impact of modifying one token depends on which others are edited simultaneously. We define an edit-set as a set of token positions to re-mask and re-sample. Motivated by prior work on sparse interactions in biological systems, we find empirically that edit-set value functions for protein inverse folding admit sparse Fourier representations. This structure enables Spectral Feedback to efficiently learn and optimize the value functions for edit-position selection. Spectral Feedback is model-agnostic and can be applied to pretrained, test-time aligned, and fine-tuned diffusion models. For all of these models, the algorithm improves alignment performance without modifying the underlying generative process. Applied to inverse folding with a protein stability reward oracle, it achieves a 32.3% increase in stable proteins for a pretrained model, 24.8% for Best-of-10, and 5.8% for a state-of-the-art RL fine-tuned diffusion model.