Look Closer: Patch-wise Supervision for AI-Generated Image Detection

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the over-reliance on global features in AI-generated image detection by exploring the efficacy of local evidence, proposing a lightweight detection framework based on explicit patch partitioning and patch-level supervision. The method shares a backbone network to independently classify cropped regions and compute losses, while achieving decision fusion during inference through probability averaging, thereby eliminating the need for residual filtering or complex fusion modules. Our investigation confirms that small-scale RGB patches retain sufficient synthetic forgery traces. Experiments on the GenImage benchmark, employing four backbone architectures spanning CNNs and Transformers, demonstrate that this patch-based strategy consistently yields significantly higher average detection accuracy than whole-image input baselines. Ultimately, this work achieves notable performance improvements through a minimalist architectural design.
📝 Abstract
How much of an image does a detector need to see? Small RGB regions can retain useful evidence of image synthesis even when they reveal little of the full scene. Motivated by single-patch detection, we study patch-wise supervision: a shared backbone classifies explicit crops, each crop receives its own loss, and patch probabilities are averaged only at inference. The procedure requires neither handcrafted residual filtering nor a learned image-level fusion module. Experiments span single-patch selection, multiple generator collections, and four CNN and Transformer backbones. On GenImage, the reported patch-wise variants improve average accuracy over their whole-image counterparts across all four backbones. Comparisons of supervision granularity, source resolution, crop size, and inference coverage further characterize the approach, while post-processing tests and difficult-image evaluation reveal its limitations. The historical experiments include evaluation-based model selection, so their scores are not presented as a uniformly selected leaderboard comparison. Overall, the study identifies explicit local input and patch-level supervision as a simple, useful combination for investigating generalizable AI-generated image detection.
Problem

Research questions and friction points this paper is trying to address.

AI-generated image detection
patch-wise supervision
local region analysis
generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Patch-wise Supervision
AI-Generated Image Detection
Local Crops
Shared Backbone
Generalizability
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhida Zhang
MAIS & NLPR, Institute of Automation, Chinese Academy of Sciences (CASIA), Beijing, China
Tao Wu
Tao Wu
ShanghaiTech University
MEMS/NEMSProcessing & EDAMultiferroicsPiezoelectric Transducers
S
Siyu Liu
Anhui University, Hefei, China
Jie Cao
Jie Cao
Institute of Automation, Chinese Academy of Sciences
Computer Vision