REALIS: A Curated Dataset for Studying the Challenges of AI Image Detection

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited generalization of AI image detectors that rely on dataset-specific cues by proposing the REALIS benchmark and evaluation framework. The dataset comprises 1.43 million real and synthetic images produced by 42 generative models. To eliminate category shortcuts, a stress-testing subset, REALIS-Expert, is constructed via semantic matching, alongside a robustness protocol covering 35 transformations. By integrating authentic prompt guidance, quality filtering, and zero-shot evaluation using multimodal large language models, the framework comprehensively assesses detector reliability under unseen generators and post-processing distribution shifts. Experimental results demonstrate that on the most challenging split, models trained on REALIS achieve a ROC-AUC of 0.752, representing a substantial improvement over the best pretrained detector, which attains only 0.550.
📝 Abstract
AI-generated image detectors are often evaluated on benchmarks where real and synthetic images differ in content, quality, or generation artifacts, allowing models to rely on dataset-specific cues and fail on unfamiliar generators or processed images. Existing datasets provide limited support for evaluating these challenges jointly across diverse visual content. We introduce REALIS, a dataset of 1.43 million real and synthetic images generated by 42 modern text-to-image models, including the latest proprietary systems such as Nano Banana 2. REALIS combines prompts derived from real images, quality filtering, and stratified sampling to reduce class-specific shortcuts while preserving content diversity. We further introduce REALIS-Expert, a stress-test subset for high-quality synthetic images, where real and generated samples are selected with closely matched semantic and visual characteristics. We also propose a robustness protocol covering 35 transformations at five severity levels to analyze detector behavior under image processing. Based on REALIS, our benchmark evaluates pretrained detectors, fine-tuned models, and zero-shot vision-language models under generator and post-processing shifts. On the hardest processed split, the best pretrained conventional detector achieves 0.550 ROC-AUC, compared with 0.752 for the best REALIS-trained detector. REALIS provides a unified framework for measuring and improving the reliability of AI-image detectors under conditions that better reflect real-world use.
Problem

Research questions and friction points this paper is trying to address.

AI image detection
dataset bias
shortcut learning
generalization
robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

AI Image Detection
Curated Dataset
Robustness Protocol
Stress-test Benchmark
Shortcut Mitigation
🔎 Similar Papers
No similar papers found.