🤖 AI Summary
This study addresses the vulnerability of Vision-Language Models (VLMs) in detecting AI-generated images, where semantic cues from embedded text can mislead authenticity judgments. To investigate this, the project systematically evaluates typographic attack strategies against VLMs by employing adversarial example construction, multimodal semantic injection, and robustness testing techniques to quantify the impact of various attack pathways on decision logic. Furthermore, it comparatively analyzes model vulnerabilities and directional asymmetries under both reasoning and direct inference modes. This work reveals a paradox wherein high detection accuracy coexists with elevated attack success rates, demonstrates that reasoning-based modes exhibit greater susceptibility to such attacks, and exposes critical security limitations in existing VLM-based detection systems when subjected to textual interference.
📝 Abstract
Vision-language models (VLMs) are increasingly used for AI-generated image (AIGI) detection, providing natural-language explanations for authenticity judgments. However, their ability to interpret text within images may also expose these judgments to misleading semantic cues. We systematically evaluate typographic attack strategies across detection-oriented, open-weight, and commercial VLMs, considering both real-to-fake and fake-to-real attacks. Our results show that reasoning modes generally exhibit greater vulnerability than direct modes and that attack effectiveness exhibits pronounced directional asymmetry. Moreover, larger models tend to exhibit higher clean detection accuracy but also higher attack success rates. We further examine attack robustness under image and text transformations and investigate whether overlays indicating the correct class can aid error correction. Together, these analyses characterize how typographic attacks influence authenticity judgments and expose limitations of current VLM-based AIGI detection systems.