🤖 AI Summary
Existing pixel-level adversarial defenses against malicious text-guided image editing attacks in diffusion models suffer from high visual detectability and weak robustness against JPEG compression and other sanitization techniques.
Method: This paper pioneers the transfer of adversarial defense to the JPEG-compatible DCT frequency domain, injecting imperceptible perturbations directly into DCT coefficients. Our approach jointly optimizes in the DCT domain, incorporates JPEG encoding-aware modeling, and enforces gradient backpropagation constraints from the diffusion-based editing model.
Results: Experiments across multiple datasets and editing tasks show near state-of-the-art editing blocking rates, significantly reduced visual distortion compared to pixel-level methods, and strong robustness against JPEG compression (quality factor ≥ 75). The proposed method thus achieves an effective trade-off between human imperceptibility and defense robustness, establishing a new paradigm for frequency-domain adversarial defense in diffusion-based vision-language editing systems.
📝 Abstract
Advancements in diffusion models have enabled effortless image editing via text prompts, raising concerns about image security. Attackers with access to user images can exploit these tools for malicious edits. Recent defenses attempt to protect images by adding a limited noise in the pixel space to disrupt the functioning of diffusion-based editing models. However, the adversarial noise added by previous methods is easily noticeable to the human eye. Moreover, most of these methods are not robust to purification techniques like JPEG compression under a feasible pixel budget. We propose a novel optimization approach that introduces adversarial perturbations directly in the frequency domain by modifying the Discrete Cosine Transform (DCT) coefficients of the input image. By leveraging the JPEG pipeline, our method generates adversarial images that effectively prevent malicious image editing. Extensive experiments across a variety of tasks and datasets demonstrate that our approach introduces fewer visual artifacts while maintaining similar levels of edit protection and robustness to noise purification techniques.