🤖 AI Summary
This study addresses the severe quality degradation of binary diffusion models under low sampling steps and their inherent training-inference mismatch. To overcome these limitations, we propose the Bernoulli Flow Model, which constructs a continuous global probability path and derives analytical posterior distributions for arbitrary time intervals. By decoupling generative dynamics from fixed discrete timesteps, this approach eliminates structural bias and enables efficient few-step self-consistent generation without distillation. Experimental results demonstrate that the proposed model achieves an FID of 9.22 on the LSUN Churches dataset with only 16 sampling steps, significantly outperforming existing baselines. These findings confirm that our method successfully combines theoretical rigor with practical effectiveness for high-quality, low-step binary image synthesis.
📝 Abstract
Binary diffusion models typically require a large number of function evaluations (NFEs) to generate high-quality samples, making practical inference computationally expensive. Reducing NFEs while preserving sample quality without distillation or additional training remains a significant challenge. Existing binary diffusion models define a discrete one-step forward path and then derive the reverse posterior. In low-NFE settings requiring cross-step sampling, they approximate the true multi-step likelihood with a single-step likelihood transition, which severely degrades sample quality. To address this fundamental limitation and decouple the generative dynamics from fixed discrete time steps, we propose Bernoulli Flow Models (BFM). Rather than relying on sequential one-step Markov diffusion chains, BFM defines a unified continuous global Bernoulli probability flow path between data distributions and pure noise, from which we derive analytical closed-form posterior transitions over arbitrary time intervals. Consequently, reducing the inference NFE is no longer an approximation based on skipping discrete steps; it only requires re-evaluating the analytical posterior over a new time grid. This eliminates the structural training-inference mismatch inherent to discrete chains and yields self-consistent low-NFE sampling. Experiments show that BFM is highly robust to aggressive NFE reduction. On LSUN Churches 256x256, a BFM trained with 256 steps achieves an FID of 9.22 using only 16 sampling steps, whereas the state-of-the-art discrete baseline degrades to 204.10. BFM also remains competitive with continuous and discrete generative baselines under standard full-step inference. These results establish BFM as a theoretically rigorous, self-consistent, and practically effective framework for fast binary data generation.