EVIE: Evidence-Vector-Informed Embeddings for Visual Document Retrieval

📅 2026-10-08
📈 Citations: 0
âœĻ Influential: 0
📄 PDF
ðŸĪ– AI Summary
This study addresses the challenges of balancing fine-grained understanding with efficient indexing in visual document retrieval, along with high OCR latency and insufficient single-vector matching accuracy. To this end, we propose a framework integrating representation learning with index compression. Methodologically, an evidence evaluation mechanism is introduced for data curation, while bidirectional symmetric teacher-student distillation combined with prefix Matryoshka representation learning optimizes model representations. Furthermore, spatially regularized hierarchical clustering is designed to achieve effective index compression. Experimental results demonstrate that the proposed approach attains state-of-the-art performance on benchmarks such as ViDoRe, substantially reducing storage overhead while significantly improving the trade-off between retrieval accuracy and storage efficiency.
📝 Abstract
Accurate and scalable visual document retrieval (VDR) requires both fine-grained page understanding and efficient indexing, yet existing approaches struggle to achieve both. OCR-based text retrieval adds preprocessing latency and can lose visual and structural cues needed to understand complex pages. Single-vector vision-language models bypass OCR, but compressing an entire page into one vector limits the granularity of query--document matching. Multi-vector retrievers with MaxSim provide finer interactions, yet demand large indexes and still leave room for accuracy improvements. We argue that overcoming these limitations requires preserving query-relevant page evidence throughout representation learning and index construction. To this end, we introduce \textbf{\textit{EVIE}} (Evidence-Vector-Informed Embeddings), a family of native visual document retrievers integrating three key innovations: (1) Evidence-judged data governance, which uses a multimodal judge to identify answer-bearing positives and filter unreliable negatives. (2) Bidirectional teacher--student learning with symmetric listwise distillation and prefix-based Matryoshka representation learning (Prefix-MRL), enabling one student checkpoint to serve six nested embedding dimensions without re-encoding. (3) Hierarchical agglomerative index compression (HAC), which clusters page tokens with spatial regularization and stores semantic centroids for single-stage MaxSim retrieval. Extensive experiments across 138 tasks from ViDoRe V1, V2, V3, and JinaVDR validate EVIE. EVIE-8B achieves 66.75 nDCG@10 on V3, exceeding the best external baseline by 1.43 points, with a four-suite average of 79.51. EVIE-4.5B with HAC retains 59.58 nDCG@10 at only 3.81 GiB per million pages, reducing vector payload by $128\times$. Together, these results improve the accuracy--storage trade-off for visual document retrieval.
Problem

Research questions and friction points this paper is trying to address.

Visual Document Retrieval
Fine-grained Page Understanding
Efficient Indexing
Multi-vector Retrieval
Accuracy-Storage Trade-off
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Document Retrieval
Evidence-Vector-Informed Embeddings
Prefix Matryoshka Representation Learning
Hierarchical Agglomerative Index Compression
Listwise Distillation
🔎 Similar Papers
No similar papers found.
💞 Related Jobs
No related jobs found.