Latent Watermarks under Generative Editing: A Benchmark and Analysis of Detection Survival

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of latent watermark detection to generative editing by establishing a systematic benchmark that evaluates the robustness of eight watermarks across five categories of editors. The methodology incorporates multiple backbone networks, varying editing intensities, and threshold recalibration strategies, while introducing the standardized clean-score separation (d') as the core evaluation metric. Experimental results reveal that editing operations effectively distinguish only the Tree-Ring watermark; notably, despite degraded score separability, detection rates remain largely sustained. Furthermore, this work validates the diagnostic efficacy of d', confirms stability at the methodological level, and demonstrates the transferability of partial findings to unseen methods. Collectively, these contributions establish a new paradigm for evaluating watermark robustness against generative manipulations.
📝 Abstract
Ordinary prompt-based editing can cause latent watermark detection to fail without explicitly targeting the watermark. We benchmark eight watermark methods against five editors across four generative backbones, four editing strengths, and five semantic categories, with edit-validity and threshold checks. Separating editing from seven subsequent distortions reveals that editing alone primarily distinguishes Tree-Ring, while added distortions expose a broader spectrum of detection survival. Sequential edits reveal a second hidden difference: score separation can decline while detection rates remain near their ceiling. Across methods, standardized clean score separation ($d'$) organizes composite-survival tiers, whereas spatial overlap adds little to predicting edit-only survival beyond clean detectability. Embedding-strength interventions in two methods link higher clean separation to higher post-edit separation. In HSTR, the margin contrast is positive, while the angular layout contrast at matched clean separation remains unresolved. Together, outcome decomposition and continuous separation expose differences hidden by aggregate TPR. Method tiers are stable under threshold recalibration at the main operating points and alternative composite weights. Clean $d'$ is thus a useful empirical diagnostic within this benchmark, with mixed transfer to unseen methods. Code and supporting artifacts are planned for a separate release.
Problem

Research questions and friction points this paper is trying to address.

Latent Watermark
Generative Editing
Detection Survival
Watermark Robustness
Benchmark
Innovation

Methods, ideas, or system contributions that make the work stand out.

Latent Watermark
Generative Editing
Benchmark
Score Separation
Detection Survival
🔎 Similar Papers
2024-08-11European Conference on Computer VisionCitations: 15
💼 Related Jobs
No related jobs found.