Ambient Discrete Diffusion: Using the Wrong Data at the Right Time for Data Efficient Learning

πŸ“… 2026-10-08
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the data scarcity challenge faced by discrete diffusion models in scientific applications by proposing RefineMix, a novel training framework. Grounded in theoretical analysis of masking mechanism properties, this method innovatively introduces out-of-distribution data exclusively during low-noise stages. By leveraging the mutual exclusivity of support sets across different domains to prevent sampling bias, RefineMix effectively balances generalization with unbiasedness. Experimental results demonstrate that RefineMix consistently outperforms baseline methods across five domain-shift settings. Furthermore, in protein sequence generation tasks, fine-tuning with merely 197 samples nearly doubles the proportion of generated novel foldable homologous proteins.
πŸ“ Abstract
We introduce RefineMix, a framework for training discrete diffusion models under severe data scarcity, a common constraint in scientific applications. RefineMix uses out-of-distribution data at selected diffusion times to improve generalization without biasing the sampling distribution. Although this strategy has been explored in continuous diffusion, discrete diffusion presents a distinct challenge: unlike Gaussian noise, masking preserves domain information in surviving tokens, limiting the use of related data at high noise levels. At low noise levels, however, the domains effectively disjoint supports become an advantage, allowing the model to learn from both in-domain and out-of-distribution data without biasing the sampler. We formalize these intuitions and provide a theoretical analysis for the proposed method. Experimentally, across five domain-shift settings, RefineMix matches or outperforms in-domain finetuning and data mixing. For protein sequence generation, finetuning with just 197 in-domain examples nearly doubles the fraction of generated proteins that are simultaneously novel, foldable, and in-family compared to standard finetuning.
Problem

Research questions and friction points this paper is trying to address.

discrete diffusion models
data scarcity
out-of-distribution data
domain shift
protein sequence generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Discrete Diffusion Models
Data Efficiency
Out-of-Distribution Data
Protein Sequence Generation
RefineMix
πŸ”Ž Similar Papers
No similar papers found.