Enhancing Document VQA Models via Retrieval-Augmented Generation

📅 2025-08-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the high memory overhead and low efficiency caused by full-page concatenation or reliance on oversized models in multi-page Document Visual Question Answering (DocVQA), this paper proposes a lightweight and efficient Retrieval-Augmented Generation (RAG) framework. Methodologically, it introduces: (i) the first OCR-free visual retrieval paradigm, synergistically combined with OCR-based text retrieval for cross-modal evidence extraction; (ii) empirical identification of the ineffectiveness of layout-aware chunking on current benchmarks, underscoring the critical role of a two-stage retrieval–reranking mechanism; and (iii) comprehensive multi-model, multi-stage evaluation demonstrating state-of-the-art gains—up to +22.5 ANLS improvement for text-based retrieval variants and +5.0 for visual variants—on MP-DocVQA, DUDE, and InfographicVQA, significantly outperforming full-page concatenation baselines.

Technology Category

Computer Vision: Image and Video RetrievalMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Question Answering

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingGraph Algorithms and Modeling for the Web: Querying, indexing, and retrieval in Web-related graphs
📝 Abstract
Document Visual Question Answering (Document VQA) must cope with documents that span dozens of pages, yet leading systems still concatenate every page or rely on very large vision-language models, both of which are memory-hungry. Retrieval-Augmented Generation (RAG) offers an attractive alternative, first retrieving a concise set of relevant segments before generating answers from this selected evidence. In this paper, we systematically evaluate the impact of incorporating RAG into Document VQA through different retrieval variants - text-based retrieval using OCR tokens and purely visual retrieval without OCR - across multiple models and benchmarks. Evaluated on the multi-page datasets MP-DocVQA, DUDE, and InfographicVQA, the text-centric variant improves the "concatenate-all-pages" baseline by up to +22.5 ANLS, while the visual variant achieves +5.0 ANLS improvement without requiring any text extraction. An ablation confirms that retrieval and reranking components drive most of the gain, whereas the layout-guided chunking strategy - proposed in several recent works to leverage page structure - fails to help on these datasets. Our experiments demonstrate that careful evidence selection consistently boosts accuracy across multiple model sizes and multi-page benchmarks, underscoring its practical value for real-world Document VQA.
Problem

Research questions and friction points this paper is trying to address.

Improving multi-page document VQA accuracy with retrieval-augmented generation
Reducing memory consumption in document VQA systems
Evaluating text-based versus visual retrieval methods for document understanding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Retrieval-Augmented Generation for document VQA
Text-based and visual retrieval methods
Evidence selection boosts accuracy significantly
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.