🤖 AI Summary
Diffusion model (DM)-generated images are increasingly indistinguishable from authentic camera-captured images, posing challenges for reliable detection.
Method: This paper proposes a lightweight, local-statistics-based detection method that addresses image spatial non-stationarity by extracting local gradient distributions and texture features. It combines handcrafted features with conventional classifiers (e.g., SVM, Random Forest), eliminating the need for deep learning training.
Contribution/Results: We systematically demonstrate—for the first time—that local statistical features significantly outperform global statistics in detecting DM-generated imagery, effectively resolving the non-stationary modeling challenge. Our method achieves substantially higher detection accuracy across diverse DM architectures compared to state-of-the-art global-statistics approaches. Moreover, it exhibits strong robustness against common post-processing operations—including JPEG compression and bilinear resizing—while maintaining high computational efficiency and practical deployability.
📝 Abstract
Diffusion models (DMs) are generative models that learn to synthesize images from Gaussian noise. DMs can be trained to do a variety of tasks such as image generation and image super-resolution. Researchers have made significant improvement in the capability of synthesizing photorealistic images in the past few years. These successes also hasten the need to address the potential misuse of synthesized images. In this paper, we highlight the effectiveness of computing local statistics, as opposed to global statistics, in distinguishing digital camera images from DM-generated images. We hypothesized that local statistics should be used to address the spatial non-stationarity problem in images. We show that our approach produced promising results and it is also robust to various perturbations such as image resizing and JPEG compression.