π€ AI Summary
Current methods for detecting AI-generated images lack reliability in high-risk scenarios and struggle to identify both synthetic content and its associated semantic anomalies. To address this, this work proposes SafeIMG, a safety-oriented benchmark encompassing 12 high-risk categories, which introduces fine-grained human annotations that jointly capture local artifacts and high-level inconsistencies with commonsense or physical lawsβenabling a shift from image-level detection toward interpretable, semantically aligned analysis. Comprehensive evaluations of specialized detectors and vision-language models (VLMs) reveal that even the best-performing VLM identifies only 49.5% of generated images (compared to 33.1% for dedicated detectors), substantially below human performance at 81.7%. Moreover, model-generated explanations cover merely 29.8% of human-annotated anomalies, with particularly poor recognition of commonsense and physical inconsistencies.
π Abstract
Rapid advances in image generation are eroding the evidentiary value of visual content in settings where authenticity can affect public safety and personal reputation. Yet existing detection benchmarks rarely examine synthetic images in public- and individual-safety contexts, where misleading visual content may carry substantial risks. Here we introduce SafeIMG, a safety-oriented benchmark spanning 12 public- and individual-safety scenarios generated using GPT Image 2. Unlike benchmarks centred on generic imagery and image-level labels, SafeIMG evaluates not only whether detectors recognise synthetic images, but also whether their decisions reflect human-identified anomalies. To this end, SafeIMG provides human annotations that localise suspicious regions and explain local artefacts and higher-level commonsense or physical inconsistencies. We evaluate specialized synthetic-image detectors and vision-language models (VLMs), and find that neither provides reliable detection. The strongest VLM identifies only 49.5% of generated images, whereas the best specialised detector identifies 33.1%, compared with 81.7% accuracy for human evaluators. Model explanations cover only 29.8\% of human-annotated anomalies and predominantly capture local defects in text, faces and hands. Their coverage falls to 15.0% for commonsense conflicts and 12.0% for physical inconsistencies, while detection performance deteriorates further after dissemination-induced image degradation. These findings show that current detectors lack the accuracy, explanatory alignment and robustness needed to evaluate AI-generated images reliably across public- and individual-safety settings.