🤖 AI Summary
This study addresses the challenge of balancing accuracy and interpretability in generative image forgery detection by proposing a weakly supervised framework grounded in visual evidence. Methodologically, it employs a modular design integrating multiple backbones, including DINOv3 and Mesorch, to extract multi-scale forensic features. A novel Grounding-DINO pseudo-mask pipeline combined with local patch-level contrastive learning generates interpretable evidence maps without requiring pixel-level annotations. Finally, a Qwen3-VL large vision-language model optimized via Group Relative Policy Optimization (GRPO) produces reasoning text. Evaluated on the XPlainVerse dataset, the proposed framework achieves a detection accuracy of 0.9349, an explanation score of 0.5571, and an overall challenge score of 0.7456, demonstrating its effectiveness in jointly delivering high detection performance and meaningful explanations.
📝 Abstract
Recent advances in generative image models have made many manipulated images highly realistic, raising the need for detectors that are not only accurate but also able to provide visual evidence for their decisions. In this paper, we present our solution to the Explainable Deepfake Detection Challenge [2] on the XPlainVerse dataset [1], where systems are required to predict whether an image is real or fake and generate both complex and simple explanations grounded in visible forensic cues. Our method follows a modular detection-and-explanation design. For the real/fake decision, we build a multi-backbone detector that combines several DINOv3 models with Mesorch manipulation-localization features, bringing together pretrained visual representations, DCT-aware cues, and multi-scale forensic information. To inject explanation evidence into the detector, we use a Grounding-DINO-based pseudo-mask generation pipeline that converts local artifact descriptions from training explanations into weak patch- level supervision for an Artifact Evidence Map. We further introduce a local patch-level contrastive objective that separates artifact and authenticity evidence in the detector feature space without requiring paired images or pixel-level manipulation masks. For language output, we use class-conditional Qwen3-VL models to generate complex explanations for fake and real predictions, followed by a GRPO-optimized text simplification model. The proposed methods were trained and evaluated on the challenge subset of XPlainVerse. On the full test split, our submission achieves 0.9349 detection accuracy, 0.5571 explanation score, and a 0.7456 final challenge score.