π€ AI Summary
This study addresses the βretrieved but failed to readβ bottleneck in document vision-language models (VLMs) by formally quantifying and introducing the retrieval-reading gap for the first time. Methodologically, a pairing protocol is designed to isolate this gap, and the FoveDoc-Bench benchmark is constructed to systematically evaluate six VLMs through the integration of OCR and visual cropping techniques. Experimental results demonstrate that injecting extracted text improves accuracy by 13β16%, confirming that textual information serves as a retrieval amplifier rather than a replacement. Furthermore, the work delineates critical boundary conditions, revealing that this approach is effective only for text-based evidence and proves ineffective or even detrimental when applied to charts and figures. The source code has been made publicly available.
π Abstract
Retrieval-augmented document question answering assumes that once the right page is found, a vision-language model (VLM) can read it. We show that this assumption often fails, leaving a retrieval-reading gap: evidence found but not used. A paired protocol isolates this gap by comparing answers from the retrieved page images alone with answers from the same images plus their extracted text. On FoveDoc-Bench, our benchmark with traceable evidence, retrieval finds nearly every evidence page, yet adding CPU-OCR text raises strict accuracy by 13 to 16 points. An exact text layer roughly doubles the gain, which appears across six VLMs from three families and, within one family, narrows with scale without closing. The reader can read this evidence but cannot find it: crops of it recover most of the text gain, boxes around it on the page none. The same protocol identifies two boundaries. Extracted text helps on textual evidence but is neutral or harmful on charts and figures. Its advantage shrinks as retrieval degrades, and unrelated text of the same form adds nothing detectable. Extracted text is an amplifier of retrieval that works, not a substitute for retrieval that does not. Our code is available at https://github.com/atoz03/fovedoc-sup.