MIC: Explaining Image-Claim Inconsistencies in AI-Generated Multimodal Misinformation

๐Ÿ“… 2026-09-27
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Existing AI-generated image detection methods focus solely on visual artifacts while neglecting image-text contextual consistency. This work proposes a multimodal inconsistency checking framework that leverages world knowledge to reveal contradictions between images and text. Methodologically, we construct the MIC-Bench benchmark and enhance the interpretability of multimodal large language models through supervised fine-tuning (SFT) combined with Group Relative Policy Optimization (GRPO), which incorporates component-level verifiable rewards. Experimental results demonstrate that the proposed approach yields significant improvements in macro-F1 scores and substantially enhances the semantic similarity of both visual evidence descriptions and world knowledge explanations. By effectively addressing the limitations of current detection paradigms, this study advances multimodal forgery detection through principled reasoning over visual-linguistic inconsistencies.
๐Ÿ“ Abstract
Claims paired with AI-generated images are a rapidly growing form of misinformation. Existing automated fact-checking (AFC) methods mainly treat this as a provenance problem, detecting low-level synthesis artifacts to decide whether an image is AI-generated. However, such methods do not verify what human fact-checkers often check: whether an image's content is consistent with the context implied by its accompanying claim. To address this gap, we introduce MIC (Multimodal Inconsistency Checking), an AFC framework that assists human fact-checkers by detecting AI-generated multimodal misinformation and explaining inconsistencies using world knowledge. MIC first uses supervised fine-tuning (SFT) for task adaptation and then applies Group Relative Policy Optimization (GRPO) to directly optimize component-level verifiable rewards for verdict prediction, inconsistency type classification, visual evidence description, and world-knowledge explanation. We further introduce MIC-Bench, a benchmark comprising 8,812 image-claim instances derived from 4,406 claims, where each claim is paired with an authentic image and an AI-generated counterpart that introduces a controlled contextual inconsistency. Compared with SFT alone, GRPO further improves Macro-F1 by 4.67 and 4.11 points in the in-distribution and out-of-distribution settings, respectively, while also improving the semantic similarity of visual evidence descriptions and world-knowledge explanations to reference annotations. Our code and data are available at https://github.com/UKPLab/arxiv2026-mic.
Problem

Research questions and friction points this paper is trying to address.

multimodal misinformation
image-claim inconsistency
automated fact-checking
AI-generated images
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Inconsistency Checking
Group Relative Policy Optimization
AI-Generated Misinformation
Automated Fact-Checking
World Knowledge
๐Ÿ”Ž Similar Papers
No similar papers found.