🤖 AI Summary
This paper investigates the fundamental differences in generative diversity among discrete latent generative models—autoregressive (AR), masked image modeling (MIM), and diffusion models. We propose the first diagnostic framework grounded in information bottleneck theory, decomposing diversity into *path diversity* (stochasticity in sampling trajectories) and *execution diversity* (output variability conditioned on a fixed trajectory), and design three zero-shot inference-time intervention methods for empirical analysis. Our findings reveal distinct trade-off strategies: MIM prioritizes diversity, AR favors compression, and diffusion enables decoupled control over path and execution diversity. Consequently, we uncover the underlying compression–diversity trade-off mechanism and introduce a plug-and-play inference-time diversity enhancement technique that significantly improves generative diversity without compromising fidelity.
📝 Abstract
Generative diversity varies significantly across discrete latent generative models such as AR, MIM, and Diffusion. We propose a diagnostic framework, grounded in Information Bottleneck (IB) theory, to analyze the underlying strategies resolving this behavior. The framework models generation as a conflict between a 'Compression Pressure' - a drive to minimize overall codebook entropy - and a 'Diversity Pressure' - a drive to maximize conditional entropy given an input. We further decompose this diversity into two primary sources: 'Path Diversity', representing the choice of high-level generative strategies, and 'Execution Diversity', the randomness in executing a chosen strategy. To make this decomposition operational, we introduce three zero-shot, inference-time interventions that directly perturb the latent generative process and reveal how models allocate and express diversity. Application of this probe-based framework to representative AR, MIM, and Diffusion systems reveals three distinct strategies: "Diversity-Prioritized" (MIM), "Compression-Prioritized" (AR), and "Decoupled" (Diffusion). Our analysis provides a principled explanation for their behavioral differences and informs a novel inference-time diversity enhancement technique.