π€ AI Summary
This work addresses the challenge of precisely localizing forged regions in AI-generated images by proposing an Iterative Forgery Amplification Network (IFA-Net), which introduces a novel detection paradigm centered on modeling the manifold of authentic images. Leveraging a frozen Masked Autoencoder (MAE) as a universal authenticity prior, IFA-Net integrates a dual-stream segmentation architecture with a task-adaptive prior injection module to establish a closed-loop iterative mechanism. This framework first identifies anomalous regions via reconstruction residuals and then converts these into guiding prompts that dynamically amplify reconstruction failure signals in suspicious areas, progressively refining pixel-level forgery localization. Evaluated on four diffusion-based image restoration benchmarks, IFA-Net achieves an average improvement of 6.5% in IoU and 8.1% in F1-score over existing methods, demonstrating superior performance and strong generalization across both unseen and conventional tampering types.
π Abstract
The proliferation of highly realistic AI-generated images poses critical challenges for digital forensics, demanding precise pixel-level localization of manipulated regions. Existing methods predominantly learn discriminative patterns of specific forgeries and often struggle with novel manipulations as editing techniques continue to evolve. We propose the Iterative Forgery Amplifier Network (IFA-Net), which shifts from learning "what is fake" to modeling "what is real". Grounded in the principle that all manipulations deviate from the natural image manifold, IFA-Net leverages a frozen Masked Autoencoder (MAE) pretrained on real images as a universal realness prior. Our framework operates through a two-stage closed-loop process: an initial Dual-Stream Segmentation Network (DSSN) fuses the original image with MAE reconstruction residuals for coarse localization, followed by a Task-Adaptive Prior Injection (TAPI) module that converts this coarse prediction into guiding prompts to steer the MAE decoder and amplify reconstruction failures in suspicious regions for precise refinement. Extensive experiments on four diffusion-based inpainting benchmarks show that IFA-Net achieves an average improvement of 6.5% in IoU and 8.1% in F1-score over the second-best method, while demonstrating strong generalization to traditional manipulation types.