🤖 AI Summary
This study addresses the limited generalization of AI image detectors that rely on dataset-specific cues by proposing the REALIS benchmark and evaluation framework. The dataset comprises 1.43 million real and synthetic images produced by 42 generative models. To eliminate category shortcuts, a stress-testing subset, REALIS-Expert, is constructed via semantic matching, alongside a robustness protocol covering 35 transformations. By integrating authentic prompt guidance, quality filtering, and zero-shot evaluation using multimodal large language models, the framework comprehensively assesses detector reliability under unseen generators and post-processing distribution shifts. Experimental results demonstrate that on the most challenging split, models trained on REALIS achieve a ROC-AUC of 0.752, representing a substantial improvement over the best pretrained detector, which attains only 0.550.
📝 Abstract
AI-generated image detectors are often evaluated on benchmarks where real and synthetic images differ in content, quality, or generation artifacts, allowing models to rely on dataset-specific cues and fail on unfamiliar generators or processed images. Existing datasets provide limited support for evaluating these challenges jointly across diverse visual content. We introduce REALIS, a dataset of 1.43 million real and synthetic images generated by 42 modern text-to-image models, including the latest proprietary systems such as Nano Banana 2. REALIS combines prompts derived from real images, quality filtering, and stratified sampling to reduce class-specific shortcuts while preserving content diversity. We further introduce REALIS-Expert, a stress-test subset for high-quality synthetic images, where real and generated samples are selected with closely matched semantic and visual characteristics. We also propose a robustness protocol covering 35 transformations at five severity levels to analyze detector behavior under image processing. Based on REALIS, our benchmark evaluates pretrained detectors, fine-tuned models, and zero-shot vision-language models under generator and post-processing shifts. On the hardest processed split, the best pretrained conventional detector achieves 0.550 ROC-AUC, compared with 0.752 for the best REALIS-trained detector. REALIS provides a unified framework for measuring and improving the reliability of AI-image detectors under conditions that better reflect real-world use.