A Bias-Free Training Paradigm for More General AI-generated Image Detection

📅 2024-12-23
🏛️ arXiv.org
📈 Citations: 1
✨ Influential: 1
📄 PDF
🤖 AI Summary
Existing AI-generated image detectors achieve strong performance on supervised benchmarks but suffer from poor generalization, primarily due to spurious correlations—such as content and formatting biases—in training data, which hinder learning of generator-specific artifacts. To address this, we propose B-Free, a bias-free training paradigm that conditions on real images and employs Stable Diffusion’s reverse sampling to synthesize semantically aligned fake counterparts, thereby establishing the first paired real-fake training framework that decouples artifact learning from content bias. Our method integrates conditional inversion, content-preserving augmentation, and bias-free contrastive learning. Evaluated across 27 diverse generative models—including FLUX and SD3.5—B-Free achieves state-of-the-art generalization, out-of-distribution robustness, and prediction calibration. Code and dataset are publicly released.

Technology Category

Computer Vision: Bias, Fairness & PrivacyMachine Learning: Adversarial Learning & RobustnessNatural Language Processing: Generation

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingSocial Networks and Social Media: Generative AI / large language models and their impact on social systemsWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Successful forensic detectors can produce excellent results in supervised learning benchmarks but struggle to transfer to real-world applications. We believe this limitation is largely due to inadequate training data quality. While most research focuses on developing new algorithms, less attention is given to training data selection, despite evidence that performance can be strongly impacted by spurious correlations such as content, format, or resolution. A well-designed forensic detector should detect generator specific artifacts rather than reflect data biases. To this end, we propose B-Free, a bias-free training paradigm, where fake images are generated from real ones using the conditioning procedure of stable diffusion models. This ensures semantic alignment between real and fake images, allowing any differences to stem solely from the subtle artifacts introduced by AI generation. Through content-based augmentation, we show significant improvements in both generalization and robustness over state-of-the-art detectors and more calibrated results across 27 different generative models, including recent releases, like FLUX and Stable Diffusion 3.5. Our findings emphasize the importance of a careful dataset design, highlighting the need for further research on this topic. Code and data are publicly available at https://grip-unina.github.io/B-Free/.
Problem

Research questions and friction points this paper is trying to address.

Detecting AI-generated images without data bias interference
Improving generalization and robustness in forensic detectors
Addressing spurious correlations in training data for better performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bias-free training using stable diffusion conditioning
Content-based augmentation for improved generalization
Semantic alignment between real and fake images
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
University Federico II | Google DeepMind
F
Fabrizio Guillaro
University Federico II of Naples
G
Giada Zingarini
University Federico II of Naples
Ben Usman
Ben Usman
Google DeepMind
Avneesh Sud
Avneesh Sud
Google DeepMind
D
D. Cozzolino
University Federico II of Naples
L
L. Verdoliva
University Federico II of Naples