Score
Designs and builds retrieval-augmented generation (RAG) systems that jointly index, retrieve, and fuse text and visual embeddings to supply multimodal context to a generative model. Implements modality-score fusion, visual-linguistic verification filters, and retrieval strategies to improve recall and relevance across homogeneous corpora.
Large language models (LLMs) suffer from knowledge staleness and hallucination due to reliance on static training data; retrieval-augmented generation (RAG) mitigates this by incorporating external, dynamic information, while multimodal RAG further integrates heterogeneous modalities—such as text, images, audio, and video—posing unique challenges in cross-modal alignment and joint reasoning. To address these, this paper proposes the first unified analytical framework for multimodal RAG, systematically organizing datasets, evaluation benchmarks, assessment dimensions, and technical pathways. It introduces a novel end-to-end taxonomy covering cross-modal retrieval, heterogeneous fusion, dynamic knowledge injection, and modality-adaptive generation. Furthermore, we open-source a standardized evaluation resource library on GitHub. This work establishes both theoretical foundations and practical paradigms for building high-fidelity, real-time-updating, controllable, and trustworthy multimodal AI systems.
To address low context relevance and hallucination risks in the image-text retrieval stage of multimodal RAG, this paper proposes a learnable dynamic relevancy scoring mechanism for re-ranking, replacing fixed top-k truncation with on-demand, adaptive selection of the most contextually relevant multimodal fragments. It introduces, for the first time, a learnable relevance score into multimodal retrieval re-ranking, jointly leveraging CLIP-based cross-modal embeddings and context-aware semantic alignment. Experiments on the COCO dataset demonstrate significant improvements: context relevance increases by 32.7%, generation accuracy rises by 26.4%, and hallucination rate decreases by 41.1%, effectively mitigating cross-modal semantic mismatch.
Existing RAG systems predominantly operate in unimodal text-only settings, struggling to handle real-world scenarios where queries and documents contain both textual and visual content. This work introduces URAG—the first unified framework for general-purpose multimodal RAG—and Nyx, a hybrid-modality retrieval model supporting joint text-image inputs and cross-modal semantic alignment. We propose an automated pipeline for high-quality multimodal question-answer data generation, yielding the NyxQA benchmark. Nyx is trained in two stages: (1) pretraining on diverse open-source multimodal data, followed by (2) supervised fine-tuning guided by visual-language model feedback signals, enabling joint optimization of retrieval accuracy and generative preference. Experiments demonstrate that Nyx maintains competitive performance on standard text-only RAG benchmarks while significantly improving answer quality in multimodal retrieval tasks—achieving both generality across modalities and practical deployability.
Existing text-based RAG systems fail to model visual structures—such as layout and figures—in multimodal documents (e.g., PDFs, scanned images), leading to information loss and suboptimal performance. To address this, we propose the first end-to-end vision-first RAG framework: it takes raw document images as input and leverages vision-language models (VLMs) to jointly perform image-level retrieval and generation, eliminating distortions introduced by OCR and text parsing. Our method innovatively unifies retrieval and generation via joint optimization and introduces cross-modal attention to enable joint layout-semantic modeling. Key components include image-embedding-based retrieval, training on synthetically augmented multimodal data, and VLM-driven fused generation. Evaluated on multimodal document question answering, our approach achieves 20–40% higher end-to-end accuracy than conventional text RAG baselines, while demonstrating superior data efficiency and cross-domain generalization.
Existing RAG research suffers from the absence of a unified, lightweight, and scalable standardized framework, hindering efficient method reproduction, comparative analysis, and evaluation. To address this, we propose RAGFlow: an open-source, modular, and highly customizable RAG research toolkit. Its key contributions are threefold: (1) a novel fine-grained modular architecture that decouples core components—including retrieval, re-ranking, prompt engineering, and evaluation—enabling flexible composition and independent optimization; (2) integrated support for 16 state-of-the-art RAG methods and 38 standardized benchmarks spanning textual and multimodal scenarios; and (3) a unified interface abstraction for LLMs and multimodal LLMs (MLLMs), coupled with efficient preprocessing and evaluation pipelines. Implemented in Python, RAGFlow significantly lowers the barrier to algorithmic experimentation and has been widely adopted for validating novel RAG methods and supporting pedagogical practice.
To address inaccurate text-to-image generation caused by knowledge inaccessibility under complex, fine-grained textual queries, this paper proposes a cross-modal sub-dimension decomposition framework that enables semantic sub-dimension–level disentangled alignment between queries and images. We introduce a novel sub-dimension retrieval augmentation paradigm, featuring a sub-query–aware hybrid sparse–dense retrieval strategy and a Pareto-optimal image set selection mechanism to support on-demand injection of multi-source visual features. Our method integrates sub-dimension–specific sparse retrieval, contrastive learning–driven cross-modal dense retrieval, and multimodal large language model–guided sub-query alignment generation. Evaluated on five benchmarks including MS-COCO, our approach significantly outperforms existing RAG-based methods: retrieval accuracy improves by 19.3%, FID decreases by 27.6%, and inference efficiency maintains linear scalability.
Existing multimodal RAG systems rely on LLMs to generate textual summaries of images and index only these summaries, leading to loss of critical visual details and contextual information—particularly detrimental to chart-text joint question answering in financial documents. This work proposes a direct multimodal embedding retrieval approach: leveraging models such as CLIP to jointly encode images and text into a shared embedding space, enabling cross-modal vector retrieval without intermediate LLM summarization and its associated information decay. Evaluated on a newly constructed financial report QA benchmark, our method achieves a 13-percentage-point absolute gain in mAP@5 (32% relative improvement) and an 11-percentage-point gain in nDCG@5 (20% relative improvement) over the text-summary baseline, significantly enhancing retrieval relevance and factual consistency of generated answers. Extensive experiments across six mainstream LLMs demonstrate the robustness and generalizability of the proposed approach.
Document understanding faces three key challenges: (1) OCR systems discard structural information; (2) multimodal large language models (MLLMs) exhibit limited contextual modeling capacity; and (3) conventional RAG frameworks struggle with heterogeneous modalities—text, tables, charts, and layout. To address these, this paper proposes a comprehensive multimodal RAG framework for document understanding. Methodologically, it introduces a novel “domain–modality–granularity” three-dimensional taxonomy and integrates OCR, MLLMs, vector retrieval, graph neural networks, and agent-based coordination to enable cross-modal joint modeling and reasoning. It further incorporates graph-structured document representation and an agent-driven retrieval–reasoning closed loop, substantially improving fine-grained comprehension, retrieval efficiency, and system robustness. The work also surveys mainstream datasets, evaluation benchmarks, and real-world applications, identifies open challenges, and provides a holistic technical roadmap for future research.
Existing RAG frameworks are largely confined to text modality, struggling to effectively retrieve and reason over real-world documents containing multimodal elements such as images, tables, and mathematical formulas. This work introduces MultiRAG—the first RAG framework supporting unified retrieval and generation across full modalities (text, image, table, formula). Its core innovation lies in constructing a dual-graph architecture: a structural graph modeling document layout and entity relations, and a semantic graph encoding cross-modal semantics; together with a hybrid retrieval mechanism integrating structural navigation and semantic matching. By representing multimodal entities explicitly and aligning them into a unified embedding space, MultiRAG enables fine-grained knowledge navigation and cross-modal semantic matching. It achieves significant improvements over state-of-the-art methods on multimodal benchmarks, especially for long-document scenarios. The code and models are publicly released, advancing RAG toward a truly multimodal paradigm.
Large vision-language models (LVLMs) suffer from weak factual consistency and poor dynamic adaptability due to static pretraining, frequent hallucinations, and the absence of external knowledge verification mechanisms. Method: This work introduces the first systematic multimodal Retrieval-Augmented Generation (RAG) framework for LVLMs. It proposes (i) cross-modal retrieval alignment, (ii) a position-bias-corrected re-ranking mechanism, (iii) retrieval-evidence-conditioned generation, and (iv) a unified self-reflective agent for dynamic evidence selection and irrelevant context suppression—all without model fine-tuning. Contribution/Results: Evaluated on multiple multimodal question answering and reasoning benchmarks, the framework achieves an average 5% performance gain, significantly improving factual accuracy and real-time external knowledge utilization while preserving model parameter integrity.