Harmful Content Generation in Text-to-Image Models: Capabilities and Moderation Limitations

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant limitations of existing safety moderation mechanisms in text-to-image models, which remain vulnerable to misuse for generating harmful content. We propose an automated prompt transformation pipeline that maps news headlines into unsafe prompts to systematically evaluate five open-source models and their community fine-tuned variants across dimensions such as violence and pornography, while quantifying the performance bottlenecks of synthetic image detectors under semantic distribution shifts. Through large-scale human evaluation, we reveal that community fine-tuned models exhibit severe vulnerability to malicious prompts, achieving an 89.2% generation rate for gory content and demonstrating critical failures in current moderation systems. This research underscores the urgent need to develop robust, large-scale defense frameworks against the abuse of generative AI technologies.
📝 Abstract
Text-to-image generative models can produce highly realistic imagery but also raise concerns about harmful misuse. While safety mechanisms exist, systematic evaluations of their effectiveness against realistic attacks remain limited. We present a systematic evaluation of harmful content generation across five open text-to-image models using an automated pipeline that transforms legitimate news captions into unsafe prompts targeting sexually explicit content, violence/gore, harmful stereotypes, self-harm, and hate speech. We evaluate both standard models with built-in safety mechanisms and community fine-tuned variants that bypass content restrictions. A human evaluation of 1,500 generated images shows high harmful-content generation rates: 89.2% for gore-related prompts, 47.6% for sexually explicit content, 43.6% for harmful stereotypes, 46.0% for hate speech, and 34.5% for self-harm, predominantly through graphic violence. Models show substantial capability for generating violent and stereotypical content, while community fine-tuned variants are particularly vulnerable to sexually explicit prompts. Generation quality is largely preserved under harmful prompting, producing imagery of sufficient fidelity to pose risks for disinformation and abuse; FLUX.1-dev produces clearly realistic harmful images in 30.9% of cases. We further evaluate automated moderation systems and find substantial detection gaps that allow unsafe images to evade filtering. Finally, we assess synthetic image detectors and show that models trained only on benign datasets perform worse on explicit content, while more diverse training data improves detection, highlighting semantic distribution gaps in current approaches. These findings expose limitations in current generation safeguards, moderation systems, and synthetic image detection, highlighting the need for stronger defenses against misuse at scale.
Problem

Research questions and friction points this paper is trying to address.

text-to-image models
harmful content generation
safety mechanisms
automated moderation
synthetic image detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

text-to-image models
automated evaluation pipeline
harmful content generation
safety moderation
synthetic image detection
🔎 Similar Papers
No similar papers found.