🤖 AI Summary
This study addresses the vulnerability of diffusion-generated images to forgery and the limited robustness of conventional pixel-level watermarking against perturbations. To this end, we propose a robust watermarking framework based on semantic-level micro-feature embedding. By constructing a semantic feature library and integrating concept programming with mask-guided editing, our method injects scene-consistent micro-features into generated images, achieving a paradigm shift from energy-constrained pixel perturbations to semantic embedding that effectively balances imperceptibility and attack resistance. Experimental results demonstrate that the proposed framework maintains exceptionally low bit error rates and high fidelity across multiple datasets under ten strong removal attacks, effectively defending against both image and video forgery.
📝 Abstract
Text-to-image diffusion models enable data-efficient "mimicry" attacks, wherein adversaries fine-tune the model on a handful of public photos to synthesize convincing forgeries of a target individual. A common countermeasure is to embed imperceptible, low-energy watermarks, yet recent studies show these signatures are brittle: modest post-processing or lightweight adversarial perturbations readily suppress detection, exposing a fundamental tension between imperceptibility and robustness. We introduce FeatMark, a watermarking framework that shifts from pixel-level, energy-starved perturbations to inconspicuous semantic features: small, scene-consistent micro-features that remain natural to humans while providing a stronger, machine-verifiable provenance signal. FeatMark builds domain-specific feature banks that encode each watermark as a compact concept program, pairing open-vocabulary semantic cues with reliable edit regions and instruction templates. It then automatically selects features that are both feasible and executable and injects them through modular, mask-guided concept editing, yielding highly localized, scene-consistent micro-edits that are difficult to perceive. We conduct extensive experiments across VGGFace2, CelebA-HQ, and WikiArt, evaluating against 10 strong watermark removal/purification attacks (including regeneration-style purification) and several bespoke adaptive attacks tailored to FeatMark, to assess perceptual fidelity, watermark detection accuracy, and robustness. We further demonstrate FeatMark's extensibility to video mimicry attacks. The results show FeatMark remains virtually impervious, withstanding all evaluated attacks with negligible bit-accuracy and fidelity degradation.