🤖 AI Summary
This study addresses the limited generalization of existing deepfake detectors to unseen manipulation methods, which stems from their difficulty in capturing transferable forensic cues. To overcome this, we propose an artifact-oriented decoupling framework that departs from the conventional forced-alignment paradigm by treating the diversity of generator-induced artifacts as complementary forensic clues for the first time. Methodologically, we introduce semantics-aligned cross-generator reconstruction coupled with a mask-based frequency-aware strategy to extract robust features, preserving artifact representations while suppressing semantic interference. Extensive experiments demonstrate that the proposed model achieves substantial performance improvements across multiple benchmark datasets in both cross-dataset and cross-generator scenarios, validating the effectiveness of our approach.
📝 Abstract
Existing image forgery detectors often suffer from generalization to unseen manipulation methods due to the limited ability to capture transferable forensic cues. Recent cross-reconstruction based methods attempt to improve generalization through semantic-artifact disentanglement, but typically align heterogeneous artifacts across generators and exclude artifact representations during reconstruction, which may overlook the inherent diversity and visual cues of manipulation artifacts. In this work, we revisit cross-reconstruction and introduce an artifact-oriented disentanglement framework for robust image forgery detection. We argue that \textbf{artifact diversity}, i.e., the intrinsic variations of manipulation artifacts introduced by different generation processes, contains complementary forensic cues rather than undesirable domain variations. Instead of enforcing explicit artifact alignment, our framework preserves diverse artifact characteristics through semantically aligned cross-generator reconstruction. Furthermore, we incorporate artifact representations into the reconstruction process and introduce a masked frequency-aware reconstruction strategy to emphasize manipulation-related residuals while reducing semantic interference. This design enables the model to learn transferable forensic representations from diverse artifacts. Extensive experiments on multiple benchmark datasets demonstrate improvements under both cross-dataset and cross-generator evaluation settings. Further analysis and ablation studies validate the effectiveness of artifact diversity preservation and artifact-aware cross-reconstruction.