Score
Designs and builds retrieval-ready multimodal corpora and indexes that combine text, images, tables, and figures; this includes extracting multimodal evidence and metadata, aligning and normalizing content across modalities, annotating modality-specific elements, and organizing the data for hybrid text–image retrieval and evaluation.
To address the limitations in flexibility and accuracy of image/video retrieval amid the explosive growth of multimodal data, this paper presents a systematic survey of Compositional Multimodal Retrieval (CMR)—a paradigm that enables precise cross-modal search by composing reference visual content (images/videos) with textual modifications. We propose the first unified taxonomy for CMR and introduce a three-tier methodological framework encompassing supervised, zero-shot, and semi-supervised paradigms: supervised approaches emphasize data augmentation, architecture design, and loss optimization; zero-shot methods leverage external knowledge-guided modality translation. The framework integrates contrastive learning, modality alignment, prompt tuning, knowledge distillation, and multi-source synthesis, and is compatible with foundational models including ViT, CLIP, and BLIP. Evaluating over 100 studies, we demonstrate consistent improvements—12–28% higher retrieval accuracy—in applications such as product search, video understanding, and person re-identification, alongside significantly enhanced generalization compared to conventional cross-modal retrieval methods.
This work addresses the limitations of existing e-commerce retrieval systems, which predominantly rely on textual information and struggle to effectively incorporate visual semantics from product images, thereby constraining cross-modal representation capabilities. To overcome this, the authors propose a two-stage alignment strategy tailored for e-commerce scenarios and introduce a novel vision-language fusion network that jointly optimizes multimodal representations of queries and products within a dual-tower architecture. By integrating domain-adaptive fine-tuning with a cross-modal alignment mechanism, the approach significantly enhances semantic complementarity between text and image modalities. Extensive experiments on a large-scale real-world e-commerce dataset demonstrate that the proposed method substantially outperforms text-only baselines and alternative multimodal fusion approaches, confirming its effectiveness and practical applicability.
To address the weak cross-modal retrieval capability of traditional RAG systems for visually rich documents (VRDs)—which contain heterogeneous elements such as text, images, tables, and charts—this paper proposes a training-free, multi-granularity joint retrieval framework. Methodologically, it introduces a novel hybrid strategy comprising hierarchical encoding, modality-aware retrieval, and layout-aware re-ranking: (1) leveraging off-the-shelf vision-language models to extract both semantic and structural features at multiple granularities; (2) designing a modality-adaptive similarity metric to enable robust cross-modal alignment; and (3) incorporating a two-stage, layout-aware re-ranking mechanism to enhance fine-grained localization. The framework operates without fine-tuning and unifies multimodal content processing. Evaluated on the MMDocIR and M2KR benchmarks, it achieves a state-of-the-art retrieval score of 65.56, demonstrating significant improvements in fine-grained recall accuracy for VRDs.
This work addresses the challenge of cross-modal retrieval among chemical reactions, molecular structures, and textual descriptions in scientific literature. We propose the first end-to-end multimodal retrieval system supporting joint queries over molecular graphs, reaction schemes, and natural language. Our method integrates chemistry-aware graph OCR for structural recognition, structured extraction of reaction information from both tabular and free-text sources, and a chemically grounded multimodal embedding framework with semantic alignment—enabling precise cross-modal matching among graph, text, and table modalities within a unified indexing architecture. The key contribution is the first full-stack, domain-specific multimodal alignment and retrieval framework for chemistry, which significantly improves accuracy and interpretability in complex reaction retrieval. Expert evaluation confirms its effectiveness in real-world research scenarios involving hybrid modality exploration.
This study addresses the cross-modal retrieval challenge for multimodal educational content—particularly computer science textbooks containing interleaved text and figures. We propose a multi-vector representation method that jointly encodes textual and visual semantics using a vision-language model (VLM), generating fine-grained multimodal embeddings and indexing them in a vector database to enable efficient cross-modal retrieval. Evaluated on over 3,600 pages of textbook material, we systematically compare four similarity metrics and find cosine similarity significantly outperforms alternatives. Benchmarking against 75 natural-language queries confirms substantial improvements in retrieval precision and practical utility within digital library settings. Our approach delivers a reproducible, scalable technical framework for intelligent discovery of multimodal educational resources, advancing the state of cross-modal semantic search in academic and pedagogical contexts.
Existing cross-modal image retrieval benchmarks lack rigorous evaluation capabilities for deep visual–linguistic co-understanding, particularly failing to support joint queries involving multi-entity images and relational text. To address this, we introduce MMIR—the first high-quality benchmark for hybrid-modality image retrieval—comprising the Entity Image (EI) dataset and the Mixed-Modality Image Retrieval (MMIR) dataset. We propose a novel “multi-entity image + relational text” query paradigm, formally defining and evaluating high-difficulty retrieval tasks that require cross-modal contextual alignment and semantic grounding. Built upon the WIT corpus, MMIR undergoes rigorous Wikidata entity alignment, human-crowdsourced validation, and cross-modal cleaning, ensuring reliability and reproducibility; it is fully open-sourced and supports both model training and strict evaluation. Empirical results demonstrate that MMIR significantly enhances evaluation validity for modeling entity associations and performing contextual reasoning in vision–language models.
Visual document retrieval faces significant challenges due to dense text, complex layouts, and fine-grained semantic dependencies, which hinder precise information access. This work presents the first systematic survey of the field in the era of multimodal large language models, establishing a comprehensive research framework that encompasses benchmark evaluation, methodological evolution, and future challenges. The study proposes a novel paradigm integrating multimodal embeddings, re-ranking models, retrieval-augmented generation (RAG), and agent-based systems to clarify the technological trajectory, identify critical bottlenecks, and offer a clear roadmap for advancing multimodal document intelligence.
This study addresses the limitation of existing keyword extraction methods, which predominantly focus on plain text while neglecting visual and audio modalities, and the absence of dedicated multimodal datasets. To bridge this gap, the authors introduce the first multimodal dataset comprising 1,000 academic papers, each annotated with human-labeled keywords and accompanied by original text, images (with OCR-extracted text), and audio transcripts (via automatic speech recognition). Leveraging this resource, the paper systematically evaluates the impact of individual modalities and their fusion—under both unsupervised and supervised settings—on keyword extraction performance. Experimental results demonstrate that integrating multimodal textual information significantly enhances extraction accuracy, with distinct modalities offering complementary cues, thereby confirming the effectiveness and necessity of multimodal modeling for academic keyword extraction.
Existing cross-modal retrieval benchmarks primarily focus on coarse-grained or single-condition alignment, falling short in addressing real-world user queries that involve multiple constraints and fine-grained specifications expressed in natural language. To bridge this gap, this work proposes MCMR—the first benchmark for multi-condition, fine-grained, and composable cross-modal retrieval—spanning five product domains and emphasizing constraint awareness and interpretability. We employ a multimodal large language model (MLLM) as both the retriever and a pointwise re-ranker, integrating visual features with long-form textual metadata for joint verification. Experiments demonstrate that visual cues dominate top-ranked accuracy, textual metadata enhances ranking stability for long-tail items, and MLLM-based re-ranking substantially improves fine-grained matching performance, thereby filling a critical evaluation gap in complex query scenarios.
This work addresses the limitations of existing vision-language models, which exhibit suboptimal performance on text retrieval tasks and incur substantial storage and inference overhead in multilingual settings due to multi-encoder architectures. To overcome these challenges, we propose a unified multitask learning framework that jointly optimizes multilingual image-to-text retrieval, text-to-image retrieval, and natural language understanding (NLU) within a single model for the first time. By sharing a common text encoder and aligning cross-modal embedding spaces with NLU-enhanced semantic representations, our approach significantly improves retrieval accuracy across modalities and languages while reducing system complexity and inference costs. This enables efficient, unified representation learning without compromising performance.