🤖 AI Summary
This study addresses the poor detection generalizability of AI-generated images caused by the absence of localized forgery traces. To this end, it proposes a detection framework based on entropy guidance and the Variational Information Bottleneck (VIB). By exploiting the inherently low-texture characteristics of generative models, the method employs image entropy metrics to select highly informative patches and integrates VIB to extract compact, discriminative representations compatible with both CNN and Transformer backbones. Experimental results demonstrate that this framework achieves 85.7% accuracy using only 2% of the training data and attains an average cross-generator accuracy of 83.5%, outperforming baseline methods by over 15%. Ultimately, this work realizes data-efficient and strongly generalizable detection of AI-generated imagery.
📝 Abstract
The proliferation of photorealistic AI-generated images demands robust detection methods that generalize across diverse generative models. While existing approaches target manipulation-based forgeries with local artifacts, generation-based images (e.g., from diffusion models) lack such traces, posing a fundamental challenge. We observe that generative models prioritize global semantics at the expense of local texture fidelity, making low-texture regions key indicators of synthetic origin. To exploit this, we propose EIB-Net, an Entropy-guided Information Bottleneck Network. EIB-Net introduces a novel Image Entropy (IE) metric to automatically select the most informative (lowest-entropy) patch, then processes it with a Variational Information Bottleneck (VIB) to learn compact, generalizable features. Extensive experiments on DIFF, DiffusionForensics, and GenImage benchmarks demonstrate state-of-the-art performance: EIB-Net achieves 85.7\% accuracy using only 2\% of training data, outperforming full-image baselines by over 15\%, and maintains robust cross-generator generalization (83.5\% average accuracy on GenImage). Furthermore, our entropy-guided patch selection (EGPL) consistently enhances diverse backbones (CNNs and Transformers), proving its practical value for data-efficient detection.