🤖 AI Summary
This study addresses the risk of scene text semantic leakage in image generation models, wherein rendered text is misinterpreted as instructions that inadvertently control non-textual regions. For the first time, this work defines rendered text as a cross-modal attack surface, decoupling visual prompts from textual semantics to reveal their dual attributes. Based on the FLUX-2-dev model, it systematically characterizes the leakage mechanism through prompt decomposition and intermediate-layer evidence analysis, further validating attack pathways augmented by large language models. The contributions include quantitatively demonstrating that harmful semantics can permeate non-textual regions and proposing a mitigation strategy that significantly suppresses unsafe semantic transfer while preserving text rendering capabilities, thereby offering novel perspectives for AI safety.
📝 Abstract
The reliability and accountability of image generative models (IGMs) are essential for building responsible and trustworthy AI systems. Recent IGMs, such as Nano Banana and GPT-Image, now support complex instruction following, realistic image synthesis, and controllable scene-text rendering. As these capabilities expand, safety analysis must also account for new control channels introduced by complex prompts. In this work, we study rendered-text semantic leakage, a largely overlooked phenomenon in open-domain text rendering. Although rendered text is intended to serve as a local visual constraint that should be reproduced verbatim in the generated image, it also carries linguistic semantics that may be interpreted by the model as part of the input instruction. This makes rendered text a potential semantic control channel whose safety implications remain insufficiently understood. We systematically characterize this phenomenon by decoupling the main visual prompt from the rendered text and measuring their individual and compositional effects on generated images. We quantify semantic leakage and rendering fidelity, and further analyze how leakage emerges from intermediate model evidence. We then show that harmful semantics embedded in scene text can persist through LLM-based prompt enhancement pipelines and steer non-text image regions, even when the main visual prompt remains benign. Finally, we propose a preliminary mitigation approach that reduces unsafe semantic transfer from rendered text to non-text regions while preserving the intended text-rendering behavior on FLUX-2-dev. Our findings reveal rendered text as a dual-use carrier of visible data and latent semantics, exposing a text-centric cross-modal attack surface in modern IGMs.