DailyBench: A Unified Benchmark for AI-Generated and Manipulated Images from Modern Generative Models

📅 2026-07-27
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current benchmarks for detecting AI-generated images predominantly rely on outdated generative models and struggle to address the challenges posed by today’s high-fidelity synthesis and localized image editing. To bridge this gap, this work introduces DailyBench, a unified benchmark comprising two subsets: FakeBench for modern full-image generation and ManipulationBench for object-level image manipulation. DailyBench is the first to jointly incorporate state-of-the-art generation and fine-grained tampering tasks, offering a high-quality testbed that closely mirrors real-world scenarios. Experimental results reveal that while leading detectors achieve 91–96% accuracy on conventional datasets like GenImage, their performance drops substantially to 54–76% on DailyBench, underscoring a critical deficiency in generalization to emerging generative techniques and subtle local manipulations.
📝 Abstract
Recent advances in generative models have shifted AI-generated image detection from identifying easily distinguishable, fully synthetic images to identifying highly realistic content generated by both modern generation and manipulation pipelines. However, existing detection benchmarks are often built with outdated generative models and primarily emphasize full-image synthesis, creating a growing mismatch between benchmark data and the images encountered in real-world generation and editing scenarios. To bridge this gap, we introduce DailyBench, a high-quality unified benchmark for evaluating whether AI-generated image detectors can generalize across both modern full-image synthesis and object-level manipulation. DailyBench contains two complementary subsets: FakeBench, which includes high-quality images synthesized by recent open-source and commercial generative models, and ManipulationBench, which introduces challenging object-level edits applied to real images using advanced image-conditional models. This design makes DailyBench a realistic testbed for studying both generator-level generalization and manipulation-aware detection under subtle local edits. Experiments on DailyBench reveal substantial robustness gaps in current detectors: methods reporting 91-96% balanced accuracy on GenImage drop to 60-76% on FakeBench and 54-66% on ManipulationBench. These results show that existing detectors remain poorly generalized to realistic synthesis and manipulation, highlighting DailyBench as a rigorous testbed for developing robust and manipulation-aware AI-generated image detection methods. The project is available at https://dailybench.github.io/
Problem

Research questions and friction points this paper is trying to address.

AI-generated image detection
generative models
image manipulation
benchmark
generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

AI-generated image detection
unified benchmark
object-level manipulation
generalization
generative models