π€ AI Summary
This study addresses the challenges of localizing geometric inconsistencies in multi-view images and the limited generalizability of existing methods. To this end, we introduce DeformView, the first wide-baseline multi-view inconsistency benchmark dataset, and propose DEFECt3R, a lightweight classifier designed for pixel-level inconsistency localization. By leveraging cross-view feature matching, DEFECt3R effectively captures inter-view geometric relationships. Furthermore, a hard negative supervision strategy is incorporated to explicitly train the model, thereby suppressing false alarms. Experimental results demonstrate that DEFECt3R significantly improves localization accuracy for geometric inconsistencies while substantially reducing the false positive rate, establishing a new benchmark for multimedia forensics.
π Abstract
Novel view synthesis (NVS) models can produce realistic new views of the same scene from different viewpoints. However, these generated views are not always geometrically consistent with one another. Multi-view (MV) consistency has shown promise as a tool for evaluating these NVS models. Its potential for multimedia forensics, however, remains largely unexplored, particularly for localizing geometric inconsistencies across wide-baseline image pairs. To enable research in this direction, we introduce DeformView, a wide-baseline MV dataset with pixel-level annotations of geometric inconsistencies. Using DeformView, we evaluate state-of-the-art MV consistency-scoring methods and show that approaches developed for NVS evaluation transfer poorly to the forensic task of geometric inconsistency localization. To address this limitation, we propose DEFECt3R, a lightweight learning-based classifier that uses cross-view feature relationships to localize geometric inconsistencies at the pixel level. By learning from explicit supervision, including hard negatives from geometrically consistent yet deformed views, DEFECt3R improves localization performance and substantially reduces false positives compared to existing consistency-scoring methods. Ablation experiments further show that both feature representations and correspondence quality contribute to localization performance. Overall, our findings demonstrate that MV geometric consistency is a promising yet underexplored signal for multimedia forensics and establish a benchmark and baseline for geometric inconsistency localization in wide-baseline MV image pairs. Code and dataset are available at https://github.com/IDLabMedia/DeformView-DEFECt3R