🤖 AI Summary
This study investigates whether multimodal large language models (MLLMs) can be exploited to generate realistic fake news and evaluates their capacity to detect such content. To this end, we propose a novel multi-agent collaborative framework in which story, image, and critic agents cooperate to synthesize over 9,000 high-quality multimodal misinformation samples, enabling a systematic assessment of the detection performance across 16 mainstream models. This work quantifies the capability boundaries of MLLMs at both ends of the adversarial spectrum: offensive generation and defensive detection. Experimental results reveal that most models exhibit detection accuracy substantially below human-level performance, with pronounced deficiencies in visual forgery identification. These findings establish an empirical foundation for developing robust defense systems against multimodal misinformation.
📝 Abstract
The rapid advancement of generative AI raises concerns about the misuse of Multimodal LLMs (MLLMs) for large-scale disinformation campaigns on social media. Despite existing research on textual disinformation, a fundamental question remains unanswered: can MLLMs be exploited to fabricate realistic multimodal fake news, and can they reliably detect it? We introduce a multi-agent framework in which a story agent, an image agent, and a critic agent collaborate to produce fake social media posts that plausibly counter true news. We apply the framework to generate over 9,000 paired multimodal news posts across science, health, and entertainment domains, and benchmark 16 open- and closed-source MLLMs for automated detection. We find that most models fall substantially short of human-level accuracy and fail critically on identifying image authenticity. Our research provides a foundation for developing robust defenses against social media fake news. Code and data are available at https: //github.com/xiuzhenzhang/Multimodal.