Keep It CALM: Analyzing the Limits of Global Unsafety in Text-to-Image Generation

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent trade-off in text-to-image models, where global safety signals struggle to simultaneously cover heterogeneous risks and preserve benign prompts. We propose CALM, a training-free method that replaces global erasure with localized counterfactual correction to precisely suppress unsafe content while retaining generative utility. The core innovation lies in exposing the dilemma of global assumptions and introducing a pioneering prompt-level adaptive local modulation mechanism based on matching anchors. This approach integrates geometric analysis, counterfactual reasoning, token representation editing, and residual component suppression to achieve precise intervention. Experimental results demonstrate that CALM significantly improves the suppression rate of unsafe content while effectively preserving generation quality for benign prompts.
📝 Abstract
Training-free safeguards for text-to-image generation often rely on a reusable safety signal, such as an unsafe direction or global toxic subspace, applied broadly across prompts. We provide a controlled geometric analysis of this global-unsafety assumption and reveal a consistent coverage-selectivity trade-off: compact unsafe subspaces fail to cover heterogeneous unsafe semantics, whereas broader aggregation increasingly distorts safety-adjacent benign prompts. Motivated by this finding, we propose CALM (Counterfactual Adaptive Local Modulation), a training-free safeguard that replaces uniform global removal with prompt-local counterfactual correction. Using matched unsafe-benign anchors, CALM routes each prompt to active unsafe categories, minimally edits only violating token representations toward the safe side, and suppresses positively aligned unsafe residual components. Across broad evaluation, CALM significantly improves unsafe content suppression while preserving benign utility, demonstrating that local counterfactual correction provides a more selective alternative to global unsafe signal removal.
Problem

Research questions and friction points this paper is trying to address.

text-to-image generation
global unsafety
training-free safeguards
coverage-selectivity trade-off
unsafe subspace
Innovation

Methods, ideas, or system contributions that make the work stand out.

Text-to-Image Generation
Training-free Safeguards
Counterfactual Adaptive Local Modulation
Global Unsafety
Safety Alignment