What Transfers from a VLM Teacher? Comparing Supervision Signals for Visual Document Retrieval
This study addresses the problem of false negative misclassification caused by incomplete annotations in visual document retrieval. We propose a hard negative discrimination and distillation framework leveraging teacher signals from Vision-Language Models (VLMs). By scoring candidate negatives with a VLM, we systematically compare knowledge transfer effects under different supervision signals and optimize the retrieval training strategy through contrastive learning and attention mechanisms. This work is the first to demonstrate that VLM-guided hard negative discrimination significantly outperforms conventional positive sample augmentation approaches. Experimental results show that the proposed method improves nDCG@5 on ViDoRe v2 to 63.0. Furthermore, we release 3.3 million teacher-annotated samples along with the source code to facilitate future research.