A Multi-Granularity Multimodal Retrieval Framework for Multimodal Document Tasks

๐Ÿ“… 2025-05-01
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
To address the weak cross-modal retrieval capability of traditional RAG systems for visually rich documents (VRDs)โ€”which contain heterogeneous elements such as text, images, tables, and chartsโ€”this paper proposes a training-free, multi-granularity joint retrieval framework. Methodologically, it introduces a novel hybrid strategy comprising hierarchical encoding, modality-aware retrieval, and layout-aware re-ranking: (1) leveraging off-the-shelf vision-language models to extract both semantic and structural features at multiple granularities; (2) designing a modality-adaptive similarity metric to enable robust cross-modal alignment; and (3) incorporating a two-stage, layout-aware re-ranking mechanism to enhance fine-grained localization. The framework operates without fine-tuning and unifies multimodal content processing. Evaluated on the MMDocIR and M2KR benchmarks, it achieves a state-of-the-art retrieval score of 65.56, demonstrating significant improvements in fine-grained recall accuracy for VRDs.

Technology Category

Computer Vision: Multi-modal VisionMachine Learning: Multimodal LearningNatural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web dataGraph Algorithms and Modeling for the Web: Querying, indexing, and retrieval in Web-related graphs
๐Ÿ“ Abstract
Retrieval-augmented generation (RAG) systems have predominantly focused on text-based retrieval, limiting their effectiveness in handling visually-rich documents that encompass text, images, tables, and charts. To bridge this gap, we propose a unified multi-granularity multimodal retrieval framework tailored for two benchmark tasks: MMDocIR and M2KR. Our approach integrates hierarchical encoding strategies, modality-aware retrieval mechanisms, and reranking modules to effectively capture and utilize the complex interdependencies between textual and visual modalities. By leveraging off-the-shelf vision-language models and implementing a training-free hybridretrieval strategy, our framework demonstrates robust performance without the need for task-specific fine-tuning. Experimental evaluations reveal that incorporating layout-aware search and reranking modules significantly enhances retrieval accuracy, achieving a top performance score of 65.56. This work underscores the potential of scalable and reproducible solutions in advancing multimodal document retrieval systems.
Problem

Research questions and friction points this paper is trying to address.

Enhancing retrieval for visually-rich documents with text and visuals
Bridging text-based retrieval limitations in multimodal document tasks
Improving accuracy via layout-aware search and reranking modules
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical encoding captures multimodal interdependencies
Modality-aware retrieval enhances visual-text integration
Training-free hybrid strategy boosts retrieval performance
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
M
Mingjun Xu
DP Technology
Z
Zehui Wang
DP Technology
Hengxing Cai
Hengxing Cai
Sun Yat-sen University
LLMVLMVLNUAV
R
Renxin Zhong
Sun Yat-sen University