MM-VeriAgent: Learning to Use Extensive Tools to Verify Multimodal Misinformation with Reinforcement Learning

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited adaptability and high inference costs of existing methods for mixed-source multimodal misinformation detection by proposing a reinforcement learning-based, tool-augmented large language model agent framework. The framework constructs a unified interface toolkit that enables the agent to adaptively invoke tools for precise verification. Furthermore, an execution caching mechanism is designed to significantly improve training efficiency while preserving multi-step reasoning capabilities. Experimental results demonstrate that the proposed approach substantially outperforms baseline models on the MMFakeBench benchmark, achieving an effective balance between detection accuracy and online inference overhead.
📝 Abstract
Real-world multimodal misinformation often involves mixed forgery sources, requiring sample-specific detection strategies. Existing tool-augmented methods rely on predefined workflows or inference-time planning, limiting adaptability or increasing inference cost. To address this issue, we introduce \textbf{MM-VeriAgent}, which learns to verify mixed-source multimodal misinformation with tools. We first build \textbf{MM-VeriTools}, a specialized toolkit for misinformation detection agents. By benchmarking various candidate models and methods on the sub-tasks required by mixed-source detection, we select the strongest for textual, visual, and cross-modal forgery analysis and encapsulate them as callable tools with a unified interface. On top of this toolkit, we train the LVLM agent with reinforcement learning to teach it how to use these tools to better solve mixed-source detection. Since many of the tools are specialized models whose online execution at every rollout severely limits RL efficiency, we further introduce \textbf{Tool-Execution Cache}, which pre-executes candidate tool calls and reuses their cached outputs during training. This preserves multi-step rollouts while reducing online tool execution, largely improving the training efficiency.Experiments on MMFakeBench demonstrate substantial accuracy gains over the base model without explicit tool search at inference time. Ablation and efficiency analyses further validate the learned tool-use policy and show that Tool-Execution Cache reduces online tool executions during training.
Problem

Research questions and friction points this paper is trying to address.

Multimodal misinformation
Mixed-source forgery detection
Tool-augmented methods
Adaptability
Inference cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning
Multimodal Misinformation Verification
Tool-Augmented Agent
Tool-Execution Cache
LVLM
Peipei Li
Peipei Li
Beijing University of Posts and Telecommunications (BUPT)
Computer VisionImage SynthesisFace Recognition
Shuhan Xia
Shuhan Xia
北京邮电大学
人工智能 多模态
S
Shengyang Liu
Beijing University of Posts and Telecommunications
Z
Zekun Li
Minzu University of China
R
Ran He
Minzu University of China; NLPR, Institute of Automation, Chinese Academy of Sciences