🤖 AI Summary
This study addresses the inherent trade-off between training data memorization and prompt-guided generation in diffusion models by proposing a training-free suppression method. Specifically, the approach redistributes cross-attention via Gaussian smoothing while enhancing the contribution of content tokens, thereby optimizing the strength of conditional allocation without requiring additional denoising evaluations to achieve alternative generation that preserves prompt consistency. Experimental results demonstrate that the proposed method attains Pareto optimality on Stable Diffusion, effectively mitigating template replication while significantly improving image preference quality.
📝 Abstract
Text-to-image diffusion models have achieved remarkable progress in image synthesis, yet can exhibit memorization by closely reproducing individual training examples. Effective mitigation must preserve useful prompt information to guide alternative depictions. We introduce a training-free method that redistributes cross-attention with Gaussian smoothing before reinforcing content-token contributions and attenuating padding contributions, without additional denoiser evaluations. With this intervention, stronger content conditioning can improve prompt alignment at comparable training-image similarity. A local analysis identifies when reinforcement preserves shared value information while redistribution reduces localized attention mass. On Stable Diffusion v1.4 and v2.0, all evaluated smoothing widths lie on the empirical Pareto frontiers for training-image similarity versus both prompt alignment and image preference. A configuration selected on Stable Diffusion reduces template reproduction in DeepFloyd IF without further tuning. These findings support jointly controlling conditioning allocation and strength to generate prompt-consistent alternatives.