π€ AI Summary
This study addresses the data scarcity challenge faced by discrete diffusion models in scientific applications by proposing RefineMix, a novel training framework. Grounded in theoretical analysis of masking mechanism properties, this method innovatively introduces out-of-distribution data exclusively during low-noise stages. By leveraging the mutual exclusivity of support sets across different domains to prevent sampling bias, RefineMix effectively balances generalization with unbiasedness. Experimental results demonstrate that RefineMix consistently outperforms baseline methods across five domain-shift settings. Furthermore, in protein sequence generation tasks, fine-tuning with merely 197 samples nearly doubles the proportion of generated novel foldable homologous proteins.
π Abstract
We introduce RefineMix, a framework for training discrete diffusion models under severe data scarcity, a common constraint in scientific applications. RefineMix uses out-of-distribution data at selected diffusion times to improve generalization without biasing the sampling distribution. Although this strategy has been explored in continuous diffusion, discrete diffusion presents a distinct challenge: unlike Gaussian noise, masking preserves domain information in surviving tokens, limiting the use of related data at high noise levels. At low noise levels, however, the domains effectively disjoint supports become an advantage, allowing the model to learn from both in-domain and out-of-distribution data without biasing the sampler. We formalize these intuitions and provide a theoretical analysis for the proposed method. Experimentally, across five domain-shift settings, RefineMix matches or outperforms in-domain finetuning and data mixing. For protein sequence generation, finetuning with just 197 in-domain examples nearly doubles the fraction of generated proteins that are simultaneously novel, foldable, and in-family compared to standard finetuning.