Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMs

📅 2026-07-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the growing challenge of pixel-level image manipulation detection in out-of-distribution (OOD) and cross-model settings, driven by the increasing photorealism of images generated by modern vision-language models (VLMs). To tackle this, we propose a concise and effective domain generalization framework that mitigates training distribution shift through balanced mini-batch sampling and incorporates a small amount of data from emerging VLMs via late-stage fine-tuning. This strategy enhances model generalization without inducing overfitting. By integrating large-scale pretraining with pixel-level localization capabilities, our method achieves substantial improvements over the current state-of-the-art approach PIXAR, yielding average gains of 26.1% in gIoU and 26.8% in cIoU on OOD VLMs such as GPT-Images-2.0 and Gemini-3.1.
📝 Abstract
Modern vision-language models (VLMs) have significantly improved image generation and editing capabilities, making pixel-level image tampering detection increasingly important yet challenging under cross-model and out-of-distribution shifts. This work studies domain generalization for pixel-level image tampering detection in modern VLMs like ChatGPT, Gemini, Qwen-Image, etc., aiming to learn tampering localization models that remain robust across diverse VLM-generated manipulation distributions. We propose a simple yet effective domain-generalized training framework built on two practical strategies. First, we introduce a balanced minibatch sampling scheme that strategically samples tampered and real images in each minibatch, preventing biased optimization toward either manipulated artifacts or clean-image priors and avoiding training collapse, ensuring that each optimization step receives proper sampled gradient signals. Second, we adopt a simple late-injection strategy, where the detector is first trained on large-scale base data until stable convergence, and then exposed to a small amount of newly selected supporting data from emerging VLM distributions, improving adaptability without overfitting to limited new domains. Together, these components provide a simple yet strong recipe for improving pixel-level tampering localization and OOD robustness across modern VLMs. Despite the conceptual simplicity, our framework outperforms the prior state-of-the-art PIXAR by a large margin of 26.1% and 26.8% relative improvement in average gIoU and cIoU, respectively, across OOD VLMs of GPT-Images-2.0, Gemini-3.1, FLUX.2, and Seedream 4.5. Our code is available at https://github.com/VILA-Lab/PIXAR-DG
Problem

Research questions and friction points this paper is trying to address.

domain generalization
image tampering detection
pixel-level localization
vision-language models
out-of-distribution robustness
Innovation

Methods, ideas, or system contributions that make the work stand out.

domain generalization
pixel-level tampering detection
vision-language models
balanced minibatch sampling
late-injection strategy