๐ค AI Summary
To address the weak cross-modal retrieval capability of traditional RAG systems for visually rich documents (VRDs)โwhich contain heterogeneous elements such as text, images, tables, and chartsโthis paper proposes a training-free, multi-granularity joint retrieval framework. Methodologically, it introduces a novel hybrid strategy comprising hierarchical encoding, modality-aware retrieval, and layout-aware re-ranking: (1) leveraging off-the-shelf vision-language models to extract both semantic and structural features at multiple granularities; (2) designing a modality-adaptive similarity metric to enable robust cross-modal alignment; and (3) incorporating a two-stage, layout-aware re-ranking mechanism to enhance fine-grained localization. The framework operates without fine-tuning and unifies multimodal content processing. Evaluated on the MMDocIR and M2KR benchmarks, it achieves a state-of-the-art retrieval score of 65.56, demonstrating significant improvements in fine-grained recall accuracy for VRDs.
๐ Abstract
Retrieval-augmented generation (RAG) systems have predominantly focused on text-based retrieval, limiting their effectiveness in handling visually-rich documents that encompass text, images, tables, and charts. To bridge this gap, we propose a unified multi-granularity multimodal retrieval framework tailored for two benchmark tasks: MMDocIR and M2KR. Our approach integrates hierarchical encoding strategies, modality-aware retrieval mechanisms, and reranking modules to effectively capture and utilize the complex interdependencies between textual and visual modalities. By leveraging off-the-shelf vision-language models and implementing a training-free hybridretrieval strategy, our framework demonstrates robust performance without the need for task-specific fine-tuning. Experimental evaluations reveal that incorporating layout-aware search and reranking modules significantly enhances retrieval accuracy, achieving a top performance score of 65.56. This work underscores the potential of scalable and reproducible solutions in advancing multimodal document retrieval systems.