π€ AI Summary
This study addresses the trade-off whereby enhanced capabilities in large language models often diminish generation diversity, exposing fundamental bottlenecks in conventional token-entropy-based decoding strategies. To overcome this limitation, we propose "Dice Planning," an inference-time paradigm that reframes diverse generation as an instruction-following problem. By incorporating an external random number generator to guide the modelβs exploration across different modes within the response space, our approach replaces token entropy with instruction following as the primary mechanism for driving diversity. This method effectively reverses the tension between model capability and generation diversity. On open-domain tasks, it significantly outperforms existing baselines, achieving a 2.4-fold improvement in the Vendi score. Furthermore, Dice Planning requires only one-tenth of the sample size to cover equivalent high-quality modes while simultaneously discovering novel ones.
π Abstract
We introduce Gacha Decoding, an inference-time method for eliciting diverse language model generations that scales with model capability. Across open-ended domains (in-the-wild chat, creative writing, planning for image generation, and protein design), Gacha Decoding significantly outperforms existing generation diversity approaches at equal quality (up to 2.4x Vendi over the next-best prior approach), reaching the same number of high-quality modes with over an order of magnitude fewer samples (11.0x) and discovering novel modes that no other approach surfaces. Our key insight is to treat diversity as an instruction-following problem: rather than relying on the LM's token entropy, we combine its instruction-following capability with randomness from an external RNG tool to scalably identify and realize distinct modes of the response space. This approach of "planning with dice" enables Gacha to invert the long-observed tension between diversity and model capability. As the underlying LM becomes a better instruction follower, diversity under Gacha Decoding consistently improves--even as its token entropy and diversity under prior approaches decline. Together, our results highlight that instruction following, rather than token entropy alone, can drive generation diversity.