construct multimodal corpora

Designs and builds retrieval-ready multimodal corpora and indexes that combine text, images, tables, and figures; this includes extracting multimodal evidence and metadata, aligning and normalizing content across modalities, annotating modality-specific elements, and organizing the data for hybrid text–image retrieval and evaluation.

constructmultimodalcorpora

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitations of existing e-commerce retrieval systems, which predominantly rely on textual information and struggle to effectively incorporate visual semantics from product images, thereby constraining cross-modal representation capabilities. To overcome this, the authors propose a two-stage alignment strategy tailored for e-commerce scenarios and introduce a novel vision-language fusion network that jointly optimizes multimodal representations of queries and products within a dual-tower architecture. By integrating domain-adaptive fine-tuning with a cross-modal alignment mechanism, the approach significantly enhances semantic complementarity between text and image modalities. Extensive experiments on a large-scale real-world e-commerce dataset demonstrate that the proposed method substantially outperforms text-only baselines and alternative multimodal fusion approaches, confirming its effectiveness and practical applicability.

e-commerceimage-text fusionmultimodal retrieval

A Multi-Granularity Multimodal Retrieval Framework for Multimodal Document Tasks

May 01, 2025
MX
Mingjun Xu
🏛️ DP Technology | Sun Yat-sen University

To address the weak cross-modal retrieval capability of traditional RAG systems for visually rich documents (VRDs)—which contain heterogeneous elements such as text, images, tables, and charts—this paper proposes a training-free, multi-granularity joint retrieval framework. Methodologically, it introduces a novel hybrid strategy comprising hierarchical encoding, modality-aware retrieval, and layout-aware re-ranking: (1) leveraging off-the-shelf vision-language models to extract both semantic and structural features at multiple granularities; (2) designing a modality-adaptive similarity metric to enable robust cross-modal alignment; and (3) incorporating a two-stage, layout-aware re-ranking mechanism to enhance fine-grained localization. The framework operates without fine-tuning and unifies multimodal content processing. Evaluated on the MMDocIR and M2KR benchmarks, it achieves a state-of-the-art retrieval score of 65.56, demonstrating significant improvements in fine-grained recall accuracy for VRDs.

Bridging text-based retrieval limitations in multimodal document tasksEnhancing retrieval for visually-rich documents with text and visualsImproving accuracy via layout-aware search and reranking modules

Multimodal Search in Chemical Documents and Reactions

Feb 24, 2025
AK
Ayush Kumar Shah
🏛️ Rochester Institute of Technology | University of Illinois Urbana-Champaign

This work addresses the challenge of cross-modal retrieval among chemical reactions, molecular structures, and textual descriptions in scientific literature. We propose the first end-to-end multimodal retrieval system supporting joint queries over molecular graphs, reaction schemes, and natural language. Our method integrates chemistry-aware graph OCR for structural recognition, structured extraction of reaction information from both tabular and free-text sources, and a chemically grounded multimodal embedding framework with semantic alignment—enabling precise cross-modal matching among graph, text, and table modalities within a unified indexing architecture. The key contribution is the first full-stack, domain-specific multimodal alignment and retrieval framework for chemistry, which significantly improves accuracy and interpretability in complex reaction retrieval. Expert evaluation confirms its effectiveness in real-world research scenarios involving hybrid modality exploration.

Combines molecular diagrams and textFacilitates retrieval of chemical reactionsSupports cross-modal linking of chemical information

Vector embedding of multi-modal texts: a tool for discovery?

Sep 09, 2025
BP
Beth Plale
🏛️ University of Oregon | Indiana University

This study addresses the cross-modal retrieval challenge for multimodal educational content—particularly computer science textbooks containing interleaved text and figures. We propose a multi-vector representation method that jointly encodes textual and visual semantics using a vision-language model (VLM), generating fine-grained multimodal embeddings and indexing them in a vector database to enable efficient cross-modal retrieval. Evaluated on over 3,600 pages of textbook material, we systematically compare four similarity metrics and find cosine similarity significantly outperforms alternatives. Benchmarking against 75 natural-language queries confirms substantial improvements in retrieval precision and practical utility within digital library settings. Our approach delivers a reproducible, scalable technical framework for intelligent discovery of multimodal educational resources, advancing the state of cross-modal semantic search in academic and pedagogical contexts.

Benchmarking retrieval performance across different similarity measuresExploring strengths and weaknesses of vision-language models for retrievalImproving discovery in multi-modal content using vector embeddings

Entity Image and Mixed-Modal Image Retrieval Datasets

Jun 02, 2025
CB
Cristian-Ioan Blaga
🏛️ Google | Microsoft

Existing cross-modal image retrieval benchmarks lack rigorous evaluation capabilities for deep visual–linguistic co-understanding, particularly failing to support joint queries involving multi-entity images and relational text. To address this, we introduce MMIR—the first high-quality benchmark for hybrid-modality image retrieval—comprising the Entity Image (EI) dataset and the Mixed-Modality Image Retrieval (MMIR) dataset. We propose a novel “multi-entity image + relational text” query paradigm, formally defining and evaluating high-difficulty retrieval tasks that require cross-modal contextual alignment and semantic grounding. Built upon the WIT corpus, MMIR undergoes rigorous Wikidata entity alignment, human-crowdsourced validation, and cross-modal cleaning, ensuring reliability and reproducibility; it is fully open-sourced and supports both model training and strict evaluation. Empirical results demonstrate that MMIR significantly enhances evaluation validity for modeling entity associations and performing contextual reasoning in vision–language models.

Introduction of new datasets for entity and mixed-modal retrievalLack of challenging benchmarks for mixed-modal image retrievalNeed for deep cross-modal contextual understanding in retrieval

Latest Papers

What's happening recently
View more

Visual document retrieval faces significant challenges due to dense text, complex layouts, and fine-grained semantic dependencies, which hinder precise information access. This work presents the first systematic survey of the field in the era of multimodal large language models, establishing a comprehensive research framework that encompasses benchmark evaluation, methodological evolution, and future challenges. The study proposes a novel paradigm integrating multimodal embeddings, re-ranking models, retrieval-augmented generation (RAG), and agent-based systems to clarify the technological trajectory, identify critical bottlenecks, and offer a clear roadmap for advancing multimodal document intelligence.

Document LayoutMultimodal Document IntelligenceMultimodal Large Language Model

This study addresses the limitation of existing keyword extraction methods, which predominantly focus on plain text while neglecting visual and audio modalities, and the absence of dedicated multimodal datasets. To bridge this gap, the authors introduce the first multimodal dataset comprising 1,000 academic papers, each annotated with human-labeled keywords and accompanied by original text, images (with OCR-extracted text), and audio transcripts (via automatic speech recognition). Leveraging this resource, the paper systematically evaluates the impact of individual modalities and their fusion—under both unsupervised and supervised settings—on keyword extraction performance. Experimental results demonstrate that integrating multimodal textual information significantly enhances extraction accuracy, with distinct modalities offering complementary cues, thereby confirming the effectiveness and necessity of multimodal modeling for academic keyword extraction.

academic papercross-modal correlationinformation richness

Existing cross-modal retrieval benchmarks primarily focus on coarse-grained or single-condition alignment, falling short in addressing real-world user queries that involve multiple constraints and fine-grained specifications expressed in natural language. To bridge this gap, this work proposes MCMR—the first benchmark for multi-condition, fine-grained, and composable cross-modal retrieval—spanning five product domains and emphasizing constraint awareness and interpretability. We employ a multimodal large language model (MLLM) as both the retriever and a pointwise re-ranker, integrating visual features with long-form textual metadata for joint verification. Experiments demonstrate that visual cues dominate top-ranked accuracy, textual metadata enhances ranking stability for long-tail items, and MLLM-based re-ranking substantially improves fine-grained matching performance, thereby filling a critical evaluation gap in complex query scenarios.

compositional matchingcross-modal alignmentfine-grained

This work addresses the limitations of existing vision-language models, which exhibit suboptimal performance on text retrieval tasks and incur substantial storage and inference overhead in multilingual settings due to multi-encoder architectures. To overcome these challenges, we propose a unified multitask learning framework that jointly optimizes multilingual image-to-text retrieval, text-to-image retrieval, and natural language understanding (NLU) within a single model for the first time. By sharing a common text encoder and aligning cross-modal embedding spaces with NLU-enhanced semantic representations, our approach significantly improves retrieval accuracy across modalities and languages while reducing system complexity and inference costs. This enables efficient, unified representation learning without compromising performance.

multilingual retrievalmultimodal retrievalretrieval efficiency

Hot Scholars

MT

Minh-Triet Tran

University of Science & John von Neumann Institute, VNU-HCM
Cryptography and SecurityMultimedia and InteractionComputer Vision and Machine LearningSoftware Engineering
HJ

Heng Ji

Professor of Computer Science, AICE Director, ASKS Director, UIUC, Amazon Scholar
Natural Language ProcessingLarge Language Models
ES

Ethan Seefried

PhD Student Colorado State University
Computer VisionVirtual RealityNatural Language ProcessingHuman Computer Interactions
KZ

Kai Zhao

Research Scientist, Walmart AI
Machine LearningArtificial Intelligence