FlowGuard: Towards Lightweight In-Generation Safety Detection for Diffusion Models via Linear Latent Decoding

📅 2026-04-09
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of detecting unsafe (NSFW) content generated by diffusion models, particularly during intermediate denoising steps where latent-space noise is difficult to monitor with existing methods. To this end, the authors propose a cross-model, in-generation safety detection framework that leverages linear latent variable decoding approximation combined with a curriculum learning strategy to identify unsafe content early in the diffusion process and terminate generation promptly. The approach establishes a unified detection benchmark spanning nine mainstream diffusion backbones and achieves F1 scores over 30% higher than current methods in both in-distribution and out-of-distribution settings. Additionally, it reduces peak GPU memory usage by 97% and accelerates latent-space projection from 8.1 seconds to just 0.2 seconds.

Technology Category

Computer Vision: Diffusion Models for VisionNatural Language Processing: Safety and RobustnessMachine Learning: Large Multimodal Models (LMMs)

Application Category

Web Mining and Content Analysis: Large pretrained models with web dataSecurity and Privacy: Large-scale security measurementsGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
Diffusion-based image generation models have advanced rapidly but pose a safety risk due to their potential to generate Not-Safe-For-Work (NSFW) content. Existing NSFW detection methods mainly operate either before or after image generation. Pre-generation methods rely on text prompts and struggle with the gap between prompt safety and image safety. Post-generation methods apply classifiers to final outputs, but they are poorly suited to intermediate noisy images. To address this, we introduce FlowGuard, a cross-model in-generation detection framework that inspects intermediate denoising steps. This is particularly challenging in latent diffusion, where early-stage noise obscures visual signals. FlowGuard employs a novel linear approximation for latent decoding and leverages a curriculum learning approach to stabilize training. By detecting unsafe content early, FlowGuard reduces unnecessary diffusion steps to cut computational costs. Our cross-model benchmark spanning nine diffusion-based backbones shows the effectiveness of FlowGuard for in-generation NSFW detection in both in-distribution and out-of-distribution settings, outperforming existing methods by over 30% in F1 score while delivering transformative efficiency gains, including slashing peak GPU memory demand by over 97% and projection time from 8.1 seconds to 0.2 seconds compared to standard VAE decoding.
Problem

Research questions and friction points this paper is trying to address.

diffusion models
NSFW detection
in-generation safety
latent diffusion
safety risk
Innovation

Methods, ideas, or system contributions that make the work stand out.

in-generation detection
latent diffusion models
linear latent decoding
NSFW detection
curriculum learning
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jinghan Yang
Fudan University
Yihe Fan
Yihe Fan
Unknown affiliation
AI safety
X
Xudong Pan
Fudan University, Shanghai Innovation Institute
Min Yang
Min Yang
Bytedance
Vision Language ModelComputer VisionVideo Understanding