Mandela-Bench: Multimodal Models Remember Canonical Images Instead of Seeing Them

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the tendency of multimodal large language models to misidentify world-knowledge-based image manipulations by over-relying on internal memory rather than visual observation. To investigate this, we propose a forgery detection paradigm based solely on knowledge verification and construct an evaluation benchmark comprising 1,507 edited images. Through anchor-controlled experiments, we conduct fine-grained fact-checking assessments across 36 mainstream models. Our findings reveal a critical failure mode wherein models prioritize parametric memory over visual evidence, yielding misjudgment rates up to 72.7%. Furthermore, we demonstrate that textual contextual prompting proves ineffective, whereas compositional cropping successfully mitigates these errors. These results expose fundamental deficiencies in the fact-verification mechanisms of existing multimodal models when confronting visually manipulated content grounded in real-world knowledge.
📝 Abstract
Historical photographs and other canonical images can now be edited seamlessly with a single instruction, often leaving no reliable pixel-level trace. In such cases, the only evidence of manipulation may be a fact about what the image depicts. Existing benchmarks instead rely on generator artefacts, image-caption inconsistencies, visual implausibilities, or external references, and therefore do not test whether a model can use its own world knowledge to verify a recognized image. We introduce Mandela-Bench, containing 1,507 edits of canonical images: 1,359 knowledge-only forgeries, each contradicting one verifiable fact, and 148 anchor-free controls that preserve the editing process without introducing a factual contradiction, together with 474 untouched originals. We score not only whether a model detects a forgery, but whether its explanation identifies the inserted entity or the fact being violated. Across 36 multimodal models, from 0.8B parameters to frontier scale, we find a consistent failure mode. When a public figure is removed from a familiar photograph, models still name that person in up to 72.7% of responses. Some models can distinguish the replacement face from the original when shown in isolation, yet still judge the full edited photograph as authentic. Providing the true event and date does not improve knowledge-grounded detection, whereas providing the same information after cropping away the recognizable composition does. Even under explicit verification prompts, only one of the 36 models meets the KGR criterion on at least half of the forged images. These results suggest that the failures cannot be explained by missing knowledge or inadequate perception alone. Instead, they are consistent with recognition biasing verification toward the remembered canonical image rather than the observed edit.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Benchmark
Knowledge-grounded Forgery Detection
Recognition Bias
Canonical Image Manipulation
World Knowledge Verification
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Y
Yicheng Bao
East China Normal University
Z
Zhenkun Gao
East China Normal University
X
Xiahui Guo
East China Normal University
M
Mingqian Yang
East China Normal University
X
Xueheng Li
University of Science and Technology of China
B
Bangwei Liu
East China Normal University
M
Mingang Chen
Shanghai Development Center of Computer Software Technology
L
Lijun Li
Shanghai Artificial Intelligence Laboratory
Xuhong Wang
Xuhong Wang
Shanghai Artificial Intelligence Laboratory
LLMKnowledge SystemAI Simulation
Xin Tan
Xin Tan
Research Professor, East China Normal University & Shanghai AI Laboratory
3D VisionTrustworthy Embodied AI