Unleashing the Potential of Vision-Language Models for Generalizable AI-Generated Image Detection

๐Ÿ“… 2026-08-05
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Existing methods for detecting AI-generated images exhibit limited generalization in complex real-world scenarios and struggle to effectively leverage the high-level semantic information from vision-language models. This work proposes Semantic Prototype Calibration (SPC), which revealsโ€”for the first timeโ€”the superior potential of Perception Encoders over mainstream vision foundation models for this task. By constructing and supervising forensic semantic class prototypes, SPC overcomes the representational limitations of conventional linear probing. Built upon a frozen Perception Encoder, the proposed method achieves new state-of-the-art performance across multiple benchmarks, substantially outperforming the DINOv3 baseline and delivering significant gains in detection accuracy, particularly in challenging in-the-wild settings.
๐Ÿ“ Abstract
Recent work has shown that a simple linear probe on frozen representations from modern vision foundation models (VFMs) can achieve state-of-the-art AIGI detection performance, substantially outperforming specialized detectors in challenging in-the-wild scenarios. This finding has established DINOv3 as the dominant foundation-model baseline for subsequent improvements. However, we find that the vision-language model Perception Encoder (PE) holds greater potential for AIGI detection, because its language-aligned representation preserves high-level provenance semantics. Specifically, PE exhibits stronger local provenance organization than DINOv3 in its frozen feature space. However, semantic-agnostic linear probing fails to exploit this structure, as PE-Linear still underperforms DINOv3-Linear by 4.1% on In-the-Wild. Based on this observation, we propose Semantic Prototype Calibration (SPC), which constructs category prototypes from forensic semantic information and calibrates them with supervised data. We apply SPC to PE and refer to the resulting detector as PE-SPC. Our analysis shows that this simple design achieves stronger generalization. Across cross-generator, post-processing, and in-the-wild benchmarks, PE-SPC surpasses the previous DINOv3 baseline and achieves new state-of-the-art results.
Problem

Research questions and friction points this paper is trying to address.

AI-generated image detection
vision-language models
generalization
provenance semantics
foundation models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
AI-Generated Image Detection
Semantic Prototype Calibration
Generalization
Perception Encoder
๐Ÿ”Ž Similar Papers
No similar papers found.
W
Weihan Cai
State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100049, China
Hao Tan
Hao Tan
Adobe Research
Vision and Language3D Multimodal
Zichang Tan
Zichang Tan
Previously CASIA, Baidu Inc.;
Computer VisionBiometricsAutonomous DrivingRoboticsMLLM
J
Jun Wan
State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences, Beijing 100190, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing 100049, China
X
Xinping Gao
Purple Mountain Laboratories, Nanjing 211111, China