Forensic Twins: Self-Supervised Residual Learning for AI-Generated Image Forensics

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited generalizability of existing AI image detectors to novel architectures and the tendency of conventional self-supervised augmentations to corrupt forensic micro-statistics. We propose Forensic Twins, a framework introducing the first self-supervised residual learning paradigm. By leveraging frozen pretrained extractors and spatially disjoint views, it suppresses macroscopic content redundancy to precisely capture acquisition pipeline fingerprints. Trained exclusively on real images without any synthetic samples, the model enables zero-shot detection. It achieves 56.61% source attribution accuracy while reducing inference latency by 375×. Furthermore, when integrated with a Gaussian Mixture Model (GMM), the approach attains 97.99% AUC across 27 unseen generators, establishing a new state-of-the-art for zero-shot AI-generated image detection.
📝 Abstract
Detectors of AI-generated images are typically trained using samples from all Generative AI architectures they must catch, and struggle as soon as a new architecture emerges. Recent approaches have explored self-supervised pre-training as an alternative solution, yet standard frameworks work against the forensic task, e.g., their augmentations overwrite the micro-statistics of image formation. This paper introduces Forensic Twins, a Self-Supervised Residual Learning (SSRL) framework whose pretext task suppresses macroscopic content availability. Each image is mapped through a frozen, off-the-shelf forensic residual extractor, from which two spatially disjoint crops are drawn. Sharing no pixel, the two views retain minimal semantic structure to align, leaving a redundancy-reduction objective with a predominant common signal: the stationary fingerprint of the image acquisition pipeline. Additionally, Forensic Twins is trained exclusively on real images; no AI-generated image is observed at any stage. Experiments show that Forensic Twins attributes AI generator sources with 56.61% accuracy, i.e., 6.13% above the previous state-of-the-art zero-shot method at 375x lower latency. We also demonstrate that fitting a Gaussian Mixture Model (GMM) offline using only the real image embeddings extracted from Forensic Twins turns it into a state-of-the-art zero-shot detector, reaching 97.99% AUC across 27 unseen AI generators, including GANs, diffusion models and commercial systems. Code, weights and exact splits will be made publicly available
Problem

Research questions and friction points this paper is trying to address.

AI-generated image forensics
zero-shot detection
self-supervised learning
generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Supervised Residual Learning
AI-Generated Image Forensics
Zero-Shot Detection
Forensic Twins
Image Acquisition Fingerprint
🔎 Similar Papers
No similar papers found.