🤖 AI Summary
Existing deep learning-based image watermarking methods rely on handcrafted noise-simulation layers, resulting in poor generalization. To address this, we propose a novel watermarking framework that eliminates the need for explicit noise-simulation training. Our key insight is the first identification and exploitation of the intrinsic robustness of image low-frequency components against both conventional signal-processing distortions and semantic editing attacks—bypassing manual distortion modeling entirely. Methodologically, we leverage a pre-trained variational autoencoder (VAE) to embed watermarks directly into the low-frequency-dominant latent feature space, tightly coupling them with structurally stable image representations. Experiments demonstrate that our approach significantly outperforms state-of-the-art methods under diverse attacks—including JPEG compression, Gaussian filtering, inpainting, and diffusion-based semantic editing—while achieving superior visual fidelity and cross-domain generalization.
📝 Abstract
The advancement of artificial intelligence generated content (AIGC) has created a pressing need for robust image watermarking that can withstand both conventional signal processing and novel semantic editing attacks. Current deep learning-based methods rely on training with hand-crafted noise simulation layers, which inherently limit their generalization to unforeseen distortions. In this work, we propose $ extbf{SimuFreeMark}$, a noise-$underline{ ext{simu}}$lation-$underline{ ext{free}}$ water$underline{ ext{mark}}$ing framework that circumvents this limitation by exploiting the inherent stability of image low-frequency components. We first systematically establish that low-frequency components exhibit significant robustness against a wide range of attacks. Building on this foundation, SimuFreeMark embeds watermarks directly into the deep feature space of the low-frequency components, leveraging a pre-trained variational autoencoder (VAE) to bind the watermark with structurally stable image representations. This design completely eliminates the need for noise simulation during training. Extensive experiments demonstrate that SimuFreeMark outperforms state-of-the-art methods across a wide range of conventional and semantic attacks, while maintaining superior visual quality.