🤖 AI Summary
This study addresses the threat of invisible watermark removal using models trained on paired data by proposing a single-image blind watermark removal method that requires neither reference images nor secret keys. The core innovation lies in a push-pull decoupling architecture, which decomposes an image via a deep encoder into structural and auxiliary residual latent variables. An adversarial push-pull supervision strategy is employed to guide watermark separation and image reconstruction, revealing the decoder's critical dependence on the auxiliary input. Experimental results demonstrate that the proposed method reduces the average bit error rate to 0.3958 while achieving a PSNR of 31.07 dB and an SSIM of 0.9554. These findings validate the effectiveness of latent space decoupling for robust watermark removal in unknown-key scenarios.
📝 Abstract
Fixed image distortions do not cover an attacker that learns from paired clean and watermarked images. We study this paired-training threat with single-image inference: deployment uses neither the clean reference nor the watermark key, payload, or decoder. An encoder maps each image to a structural latent $g$ and an auxiliary residual latent $u$. Push supervision reconstructs the watermarked image from $D(g_w,u_w)$. Pull supervision trains the zero-auxiliary output $D(A_g(g_w;k),0)$ toward the paired clean image. At $k=1.10,u=0$, the four-method sweep gives an average BER of $0.3958$, PSNR of $31.07$ dB, and SSIM of $0.9554$. Restoring $u$ from $0$ to $0.15$ moves average BER from $0.3893$ to $0.3357$, while PSNR falls from $30.99$ to $28.23$ dB. The intervention supports decoder dependence on the auxiliary input in the evaluated setting. The accompanying theory is a conditional, post-hoc account of this behavior rather than an experimentally verified information-relocation result.