🤖 AI Summary
This study addresses the vulnerability of text-to-image models in generating unsafe content and the limitations of existing pre- and post-processing safeguards against adversarial attacks and emergent unsafe signals. We propose a real-time, intra-generation safety framework that introduces a novel "mid-generation detection with controllable guidance" mechanism. Specifically, our approach monitors the denoising trajectories of diffusion models and extracts intermediate-layer features to identify NSFW signals, subsequently employing reinforcement learning to dynamically steer unsafe generation paths toward safe latent spaces. Experimental results demonstrate that the proposed method significantly outperforms existing baselines on both standard and adversarial benchmarks. It effectively suppresses policy-violating outputs while preserving high perceptual quality and prompt fidelity.
📝 Abstract
Recent Text-to-Image (T2I) models achieve remarkable visual image generation performance, but they can still generate NSFW (Not-Safe-For-Work) contents, including violent or explicit images. Existing safety checker mechanisms are largely confined to pre-generation filtering (e.g. prompt-level text classifiers) or post-hoc moderation applied after an image is completely synthesized. However, adversarial attack methods operate over a much broader space. This imbalance highlights the need for a safety mechanism that intervenes during the generation process. We propose an in-generation safety framework that monitors the denoising trajectory and detects emerging NSFW signals from intermediate representations. Rather than merely detecting NSFW generations, our method applies reinforcement learning to generate safe images from NSFW prompts. By coupling in-generation detection with controllable steering, our approach mitigates unsafe trajectories even when NSFW signals emerge after generation has already begun. Experiments results show that our method consistently outperforms existing safe image generation methods across both standard and adversarial evaluation sets, while preserving perceptual quality and prompt fidelity. Code will be released upon acceptance.