Robust Watermarks Meet Backdoored Models: Evading Diffusion Semantic Watermarks via Stealthy Backdoor

📅 2026-08-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the vulnerability of semantic watermark detection mechanisms in neural diffusion models to backdoor attacks. It systematically uncovers this security flaw for the first time and proposes GhostVAE, a novel method that implants a highly stealthy backdoor into the encoder of a variational autoencoder. GhostVAE employs a two-stage strategy to effectively evade detection: it constructs a universal trigger via power spectrum regularization and trains the backdoored encoder with a parameter alignment objective. The approach is compatible with diverse diffusion models and watermarking schemes. Extensive experiments demonstrate that GhostVAE achieves an average true positive rate of 94.4% and an attack success rate of 94.6% across three mainstream watermarking frameworks and diffusion models, while successfully bypassing 17 representative defense mechanisms.
📝 Abstract
Although semantic watermarking is considered a promising safeguard for images generated by Latent Diffusion Models (LDMs), the reliance of the watermark detection pipeline on neural networks introduces a critical yet underexplored backdoor attack surface. To systematically study this vulnerability, we propose GhostVAE to plant a stealthy backdoor into the encoder of Variational Autoencoder (VAE), enabling reliable evasion of watermark detection. GhostVAE operates in two stages: it first constructs a universal trigger via power spectrum regularization to improve the trigger robustness, and then trains a backdoored VAE encoder with a parameter-aligned objective. Through extensive evaluations across three state-of-the-art semantic watermarking schemes and three widely adopted LDMs, we show that GhostVAE preserves watermark detection performance on benign images (achieving an average true positive rate of 94.4%), while simultaneously enabling highly effective evasion under trigger activation (achieving an average attack success rate of 94.6%). Moreover, we comprehensively analyze seventeen representative defenses and demonstrate that GhostVAE remains stealthy across the input space, parameter space, and latent space. Our work fundamentally undermines the trustworthiness of semantic watermarking systems and highlights that secure deployment of semantic watermarks requires end-to-end security considerations, particularly for neural network components.
Problem

Research questions and friction points this paper is trying to address.

semantic watermarking
backdoor attack
Latent Diffusion Models
watermark evasion
VAE encoder
Innovation

Methods, ideas, or system contributions that make the work stand out.

semantic watermarking
backdoor attack
Latent Diffusion Models
GhostVAE
stealthy evasion