🤖 AI Summary
This study addresses the performance bottleneck in image forgery localization caused by the implicit modeling of artifacts, reformulating the task as a latent variable problem and proposing a two-stage paradigm to explicitly model tampering artifacts. Methodologically, it introduces paired artifact learning alongside standard localization strategies, and designs an edit-relation-based feature disentanglement mechanism that effectively separates content from artifact representations. Additionally, a large-scale dataset, EditGroup-45K, is constructed. Experimental results demonstrate that the proposed approach not only significantly enhances the localization performance of various mainstream models but also thoroughly validates its capability to explicitly capture tampering artifacts.
📝 Abstract
Image Manipulation Localization (IML) is commonly formulated as a fully supervised learning task that estimates the optimal manipulation mask $y$ for a given image $x$. In this work, we first reveal the latent nature of artifacts and thus reinterpret IML as a latent-variable problem, $P(y|x)=\int P(y|z)\,P(z|x)\,dz$, where $z$ denotes the artifacts. Following this interpretation, we pinpoint the cause for the current IML models' insufficiency as their implicit artifacts modeling strategy, highlighting the necessity of modeling $z$ in an explicit manner. Without direct labels, feature disentanglement is the most appropriate solution for this explicit modeling. Accordingly, we propose a two-stage learning paradigm with the Pairwise Artifacts Learning (PAL) and Standard Localization (SL) phases to estimate $P(z|x)$ and $P(y|z)$ via edit relations. To support our edit-relation-based learning, we further curate EditGroup-45K, a source-anchored dataset organized into edit groups for pair construction. Extensive experiments show that our PAL paradigm yields consistent improvements across diverse IML architectures, and empirical analyses further verify that PAL does capture artifacts explicitly through feature disentanglement. Code and dataset are available at https://github.com/venus-guangjian/PAL