Safe Image Generation via Reinforcement Learning

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of text-to-image models in generating unsafe content and the limitations of existing pre- and post-processing safeguards against adversarial attacks and emergent unsafe signals. We propose a real-time, intra-generation safety framework that introduces a novel "mid-generation detection with controllable guidance" mechanism. Specifically, our approach monitors the denoising trajectories of diffusion models and extracts intermediate-layer features to identify NSFW signals, subsequently employing reinforcement learning to dynamically steer unsafe generation paths toward safe latent spaces. Experimental results demonstrate that the proposed method significantly outperforms existing baselines on both standard and adversarial benchmarks. It effectively suppresses policy-violating outputs while preserving high perceptual quality and prompt fidelity.
📝 Abstract
Recent Text-to-Image (T2I) models achieve remarkable visual image generation performance, but they can still generate NSFW (Not-Safe-For-Work) contents, including violent or explicit images. Existing safety checker mechanisms are largely confined to pre-generation filtering (e.g. prompt-level text classifiers) or post-hoc moderation applied after an image is completely synthesized. However, adversarial attack methods operate over a much broader space. This imbalance highlights the need for a safety mechanism that intervenes during the generation process. We propose an in-generation safety framework that monitors the denoising trajectory and detects emerging NSFW signals from intermediate representations. Rather than merely detecting NSFW generations, our method applies reinforcement learning to generate safe images from NSFW prompts. By coupling in-generation detection with controllable steering, our approach mitigates unsafe trajectories even when NSFW signals emerge after generation has already begun. Experiments results show that our method consistently outperforms existing safe image generation methods across both standard and adversarial evaluation sets, while preserving perceptual quality and prompt fidelity. Code will be released upon acceptance.
Problem

Research questions and friction points this paper is trying to address.

Text-to-Image generation
NSFW content
safety mechanism
adversarial attacks
in-generation safety
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
In-generation Safety
Text-to-Image Generation
Denoising Trajectory Monitoring
NSFW Mitigation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
E
Eungyeol Han
Yonsei University, School of Integrated Technology
Jong-Seok Lee
Jong-Seok Lee
Professor at Yonsei University
multimedia processingmachine learning