🤖 AI Summary
Existing image forgery detection methods lack interpretable evidence or rely solely on post-hoc explanations. This work proposes ATAR, a framework that emulates the workflow of forensic experts by integrating 22 forensic tools to achieve autonomous detection, localization, and interpretation. Its core innovation lies in introducing the first dual-stream forensic reasoning paradigm coupled with a curriculum learning mechanism. By combining multimodal large language model agents with reinforcement learning, ATAR constructs closed-loop reasoning from semantic anomalies to objective evidence. Evaluated on zero-shot benchmarks, ATAR achieves an average F1 score of 78.5%, surpassing the strongest baseline by 11.8 percentage points while generating substantially more credible interpretations.
📝 Abstract
Conventional image forgery detection methods produce binary scores or pixel-level masks without interpretable evidence, while recent multimodal large language model (MLLM)-based approaches generate post-hoc explanations of predetermined classification results rather than reasoning from evidence. Inspired by the forensic workflow of human judicial experts, we propose Agentic Tool-Augmented Reasoning (ATAR), a framework integrating 22 specialized forensic tools across seven complementary domains to autonomously detect, localize, and explain image forgeries through multi-turn reasoning. A Dual-Stream Forensic Reasoning paradigm combines a high-level semantic anomaly path, which magnifies suspicious regions for fine-grained inspection, with a low-level forgery artifact path, which invokes forensic tools to extract objective evidence. We further introduce Forensics Curriculum Learning: during General Experience SFT, an automated teacher-student mentoring pipeline synthesizes multi-turn tool-usage reasoning trajectories; during Forensic Scene RL, a Tool Prior Curriculum guides early tool exploration and progressively transfers control to the agent, while a Structured Evidence Reward provides fine-grained process-level supervision. Experiments across IMDL, Deepfake detection, DMDL, and AIGC detection show that ATAR achieves 78.5% average image-level F1 on six zero-shot IMDL benchmarks, surpassing the strongest MLLM baseline by 11.8 percentage points, and remains competitive with specialized detectors on other tasks while producing substantially more faithful and grounded explanations.