🤖 AI Summary
This study addresses the inherent trade-off between satisfaction and diversity in text-to-image generation, where existing methods often sacrifice quality for diversity. To this end, we propose SatisDive, a training-free inference framework that pioneers formulating the generation process as a satisficing problem. Specifically, it enforces a reward lower bound to guarantee individual image quality while employing batch-level relative reward truncation to dynamically filter candidate images, thereby enhancing batch diversity. Extensive experiments on the FLUX.1-dev and SANA models using the Pick-a-Pic benchmark demonstrate that SatisDive improves the worst-case candidate reward by up to 0.70 compared to FK steering. Furthermore, its satisfaction-diversity curves comprehensively Pareto-dominate those of existing approaches, achieving an optimal balance between generation quality and diversity.
📝 Abstract
Text-to-image generation enables users to explore several images generated from the same prompt. For these generated images to be useful, each one must reflect the user's preferences, measured by a learned reward, and differ visually from the others to maintain diversity. Existing methods are limited: they either address reward and diversity separately or combine them in one aggregate score, enabling high diversity to offset low rewards. In this paper, we address these limitations by formulating generation as satisficing: every image (candidate) must satisfy a reward floor and the batch of images must satisfy a diversity cutoff. The reward floor controls the balance between worst-candidate reward and batch diversity; we show that varying this floor defines a Pareto frontier. To traverse this frontier, we introduce SatisDive, a training-free inference-time method. SatisDive uses a batch-relative reward cutoff to distinguish lower- from higher-reward candidates, emphasizing reward improvement for candidates below the cutoff and diversity among candidates above it. On Pick-a-Pic, at matched DreamSim, SatisDive improves worst-candidate reward over FK steering by up to 0.43 with FLUX.1-dev as the base model and HPSv3 as the reward, and by up to 0.70 with SANA-1.6B as the base model and ImageReward as the reward. More broadly, across their overlapping DreamSim ranges, SatisDive's satisfaction-diversity curve Pareto-dominates FK steering's curve in each setting.