🤖 AI Summary
This study addresses the challenge that existing text-to-image models struggle to control visual attention allocation and lack methods for enhancing target saliency without requiring visual priors. To this end, this work pioneers the task of target saliency enhancement by proposing the GazeME framework. Motivated by insights into relative saliency, the framework introduces a lightweight learnable token insertion mechanism. Through saliency token prompting and a Saliency Prior Marker Activation strategy, it precisely modulates object saliency in a prior-free manner. Furthermore, a dedicated dataset is constructed to validate the proposed approach. Experimental results demonstrate that GazeME effectively enhances the visual prominence of target objects while preserving text-semantic alignment and overall image generation quality.
📝 Abstract
Text-to-image generation has advanced in controlling what, where, and how objects appear, yet how visual attention is distributed among objects remains largely unexplored. In this paper, we introduce Target Saliency Boosting, a new task aimed at boosting the visual saliency of a specific object during text-to-image generation without requiring any visual priors. Our key insight is that visual saliency is inherently relative: boosting the saliency of a target object also depends on the global saliency distribution across all objects in the scene. Based on this insight, we propose GazeME, a lightweight framework that uses saliency-marked prompts, inserting learnable marker tokens around object descriptions to indicate which objects to visually emphasize or suppress. To learn these markers, we construct a saliency-semantics dataset that associates objects in image--prompt pairs with object-level saliency scores, and propose Saliency Prior Marker Activation (SPMA), a saliency-aware stochastic marker activation strategy that exploits relative saliency relationships for robust training. During inference, GazeME automatically inserts appropriate markers into the prompt, thereby directly enhancing the visual saliency of the target object. Extensive experiments demonstrate that GazeME effectively boosts target saliency while preserving both semantic alignment and image quality.