D-DOIT: Training-free Adaptation of Discrete Diffusion via Doob's h-Transform

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of achieving efficient, training-free adaptation of discrete diffusion models under general reward constraints. To this end, we propose the D-DOIT framework, which for the first time derives the Doob h-transform in the discrete domain to reconstruct reverse transition kernels, thereby enabling gradient-free, reward-guided sampling. Furthermore, a prediction head is introduced to approximate costly future rollouts, complemented by a late-stage Best-of-K refinement strategy to enhance generation quality. Extensive evaluations on DNA design and protein inverse folding tasks demonstrate that D-DOIT significantly outperforms existing baselines, yielding substantial improvements in sequence activity, specificity, and success rate. This work establishes an efficient new paradigm for reward alignment in discrete diffusion models.
📝 Abstract
We propose D-DOIT (Discrete Doob-Oriented Inference-time Transformation), a training-free and efficient adaptation method for discrete diffusion models with generic rewards. D-DOIT formulates adaptation as sampling from a reward-tilted target distribution and realizes this transport through Doob's h-transform of the discrete diffusion reverse kernel, using only reward values rather than reward gradients. Unlike continuous diffusion, masked discrete diffusion samples categorical token-reveal transitions rather than continuous state updates. D-DOIT derives the corresponding discrete Doob's h-transform, which guides sampling by reweighting reverse transition probabilities instead of adding a drift correction. To make this transformation practical, D-DOIT avoids expensive future rollouts. At each guided step, D-DOIT samples candidate next states, uses the model prediction head to complete each candidate into a clean sequence, evaluates each completion with the reward oracle, and resamples the next state with probabilities proportional to the rewards. An optional late-stage best-of-K refinement further improves sample quality by branching trajectories only near the end of denoising, avoiding the $K$-fold cost over the full trajectory. Empirically, across regulatory DNA design and protein inverse folding benchmarks, D-DOIT outperforms training-free guidance baselines. It improves enhancer activity and cell-type specificity while preserving sequence naturalness, and achieves the highest success rate in protein inverse folding.
Problem

Research questions and friction points this paper is trying to address.

discrete diffusion models
training-free adaptation
reward-guided generation
Doob's h-transform
protein inverse folding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Discrete Diffusion Models
Doob's h-Transform
Training-free Adaptation
Reward-guided Sampling
Best-of-K Refinement
🔎 Similar Papers
J
Jieke Wu
Department of Computer Science, King Abdullah University of Science and Technology
Q
Qijie Zhu
Department of Statistics and Data Science, Northwestern University
Weimin Wu
Weimin Wu
Ph.D. Candidate in Computer Science, Northwestern University
AI for BiologyML Theory
Z
Zeqi Ye
Industrial Engineering & Management Sciences, Northwestern University
Minshuo Chen
Minshuo Chen
Northwestern University
Diffusion ModelReinforcement LearningGenerative Modeling
H
Han Liu
Center for Foundation Models and Generative AI, Northwestern University; Department of Statistics and Data Science, Northwestern University