multimodal retrieval systems

Designs and implements systems that retrieve and rank visual and textual content by semantic similarity across modalities—covering content-based image and video retrieval, image–text matching, multimodal candidate generation and ANN indexing, and exemplar or similar-case lookup. Work includes building multimodal embedding spaces and cross-modal encoders, search indexes and rerankers, relevance metrics and bias-aware ranking strategies, and integrating retrieved exemplars for downstream retrieval-based reasoning.

multimodalretrievalsystems

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.33
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitations of existing e-commerce retrieval systems, which predominantly rely on textual information and struggle to effectively incorporate visual semantics from product images, thereby constraining cross-modal representation capabilities. To overcome this, the authors propose a two-stage alignment strategy tailored for e-commerce scenarios and introduce a novel vision-language fusion network that jointly optimizes multimodal representations of queries and products within a dual-tower architecture. By integrating domain-adaptive fine-tuning with a cross-modal alignment mechanism, the approach significantly enhances semantic complementarity between text and image modalities. Extensive experiments on a large-scale real-world e-commerce dataset demonstrate that the proposed method substantially outperforms text-only baselines and alternative multimodal fusion approaches, confirming its effectiveness and practical applicability.

e-commerceimage-text fusionmultimodal retrieval

To address the limitations in flexibility and accuracy of image/video retrieval amid the explosive growth of multimodal data, this paper presents a systematic survey of Compositional Multimodal Retrieval (CMR)—a paradigm that enables precise cross-modal search by composing reference visual content (images/videos) with textual modifications. We propose the first unified taxonomy for CMR and introduce a three-tier methodological framework encompassing supervised, zero-shot, and semi-supervised paradigms: supervised approaches emphasize data augmentation, architecture design, and loss optimization; zero-shot methods leverage external knowledge-guided modality translation. The framework integrates contrastive learning, modality alignment, prompt tuning, knowledge distillation, and multi-source synthesis, and is compatible with foundational models including ViT, CLIP, and BLIP. Evaluating over 100 studies, we demonstrate consistent improvements—12–28% higher retrieval accuracy—in applications such as product search, video understanding, and person re-identification, alongside significantly enhanced generalization compared to conventional cross-modal retrieval methods.

Address challenges in supervised and zero-shot learning paradigmsEfficiently search heterogeneous multi-modal dataEnhance search flexibility with composed multi-modal retrieval

This paper addresses the modality-asymmetric retrieval problem in e-commerce search, where queries are purely textual while items are multimodal (text-image). To tackle cross-modal representation fusion and semantic alignment, we propose SMAR—a Semantic-enhanced Multimodal Alignment and Retrieval model. SMAR introduces a novel cross-modal alignment mechanism: it models fine-grained text-image associations via deep semantic matching, employs an attention-driven modality interaction module for dynamic feature alignment, and jointly optimizes the text and image encoders in an end-to-end manner. Evaluated on a large-scale industrial e-commerce dataset, SMAR significantly outperforms state-of-the-art unimodal and multimodal baselines, achieving an average +8.2% improvement in Recall@10. Furthermore, we release the first large-scale, publicly available triplet dataset—comprising e-commerce queries, product images, and corresponding textual descriptions—specifically designed for asymmetric multimodal retrieval, thereby enabling reproducible research in this emerging direction.

Addressing multimodal retrieval with visual and textual dataEnhancing e-commerce search with semantic retrievalSolving modality fusion in asymmetric query-item scenarios

Multimodal Representation Alignment for Cross-modal Information Retrieval

Jun 10, 2025
FX
Fan Xu
🏛️ University of Luxembourg

Cross-modal retrieval suffers from semantic misalignment due to geometric inconsistency between image and text representations. This work formulates image–text matching as a metric alignment problem in the embedding space. We systematically demonstrate that: (1) the Wasserstein distance quantitatively characterizes inter-modal distributional discrepancy; (2) cosine similarity exhibits superior robustness over Euclidean distance, KL divergence, and other conventional metrics for alignment; and (3) standard MLPs inadequately capture complex cross-modal interactions. To address these issues, we propose a lightweight neural metric learning framework that integrates vision-language models with unimodal encoders, jointly optimizing four classical metrics and two learnable neural metrics. Our approach achieves significant improvements in Recall@K across standard cross-modal retrieval benchmarks. Moreover, it provides practical, deployment-oriented alignment evaluation criteria and architectural design guidelines.

Align multimodal representations for cross-modal retrievalImprove feature alignment with cosine similarityMeasure modality gap using Wasserstein distance

Real-world documents—such as PDFs, slides, and videos—contain rich heterogeneous visual and semantic information; conventional text-based retrievers suffer from performance bottlenecks due to their reliance on structured textual inputs. To address this, we propose the first unified multimodal retrieval framework supporting end-to-end joint modeling of text, images, audio, and video—breaking the constraints of unimodal paradigms. Our method leverages state-of-the-art multimodal large language models (e.g., Qwen2.5-Omni) to generate image-augmented document representations, enabling deep cross-modal alignment and learning of a shared embedding space for both cross-modal and fused-modal retrieval. Evaluated on diverse multimodal benchmarks, our approach significantly outperforms existing methods, demonstrating strong generalization to unstructured documents and practical applicability in real-world scenarios. The core contributions are: (1) the first end-to-end four-modal unified retrieval architecture, and (2) systematic improvements in multimodal content understanding and retrieval accuracy.

Enabling cross-modal and joint-modal retrieval using single modelOvercoming limitations of text-based retrievers with rich contentUnified multimodal retrieval for text, image, audio, and video

Latest Papers

What's happening recently
View more

This work investigates the potential of multimodal large language models (MLLMs) to perform purely visual tasks without any training, with a focus on instance-level similarity assessment in large-scale image retrieval. The authors propose a zero-shot reranking method that feeds image pairs into an MLLM and converts its next-token prediction probabilities into similarity scores. Coupled with a memory-efficient indexing mechanism, this approach enables scalable top-k reranking. Notably, it is the first to directly apply MLLMs—without fine-tuning or task-specific architectures—to training-agnostic large-scale image retrieval reranking. The method outperforms specialized rerankers trained on non-native domains across multiple benchmarks and demonstrates superior robustness in challenging scenarios such as cluttered backgrounds, occlusions, and small objects.

Image RetrievalInstance-level SimilarityLarge-scale Vision Tasks

Existing cross-modal retrieval benchmarks primarily focus on coarse-grained or single-condition alignment, falling short in addressing real-world user queries that involve multiple constraints and fine-grained specifications expressed in natural language. To bridge this gap, this work proposes MCMR—the first benchmark for multi-condition, fine-grained, and composable cross-modal retrieval—spanning five product domains and emphasizing constraint awareness and interpretability. We employ a multimodal large language model (MLLM) as both the retriever and a pointwise re-ranker, integrating visual features with long-form textual metadata for joint verification. Experiments demonstrate that visual cues dominate top-ranked accuracy, textual metadata enhances ranking stability for long-tail items, and MLLM-based re-ranking substantially improves fine-grained matching performance, thereby filling a critical evaluation gap in complex query scenarios.

compositional matchingcross-modal alignmentfine-grained

This work addresses the limitation of existing large models in cross-modal retrieval, which often neglect subject-level semantics, leading to visual oversight and semantic drift that hinder precise alignment between key image regions and textual descriptions. To overcome this, the authors propose the SSA-ME framework, which introduces, for the first time, a subject-level saliency modeling mechanism. This mechanism employs saliency-aware guidance to direct cross-modal attention toward semantic cores and integrates a feature regeneration module to recalibrate visual features, thereby achieving balanced and semantically consistent fusion across modalities. Evaluated on the MMEB benchmark, the proposed method achieves state-of-the-art performance, significantly enhancing fine-grained retrieval accuracy while offering strong interpretability.

Cross-Modal RetrievalMultimodal EmbeddingSemantic Drift

Existing image-text retrieval benchmarks struggle to evaluate models’ capabilities in domain-specific knowledge and complex multimodal reasoning. To address this gap, this work proposes the first multimodal retrieval benchmark structured along two axes: knowledge depth (spanning 5 major categories and 17 subcategories) and reasoning complexity (encompassing 6 types). The benchmark includes 16 visual data types and supports both multimodal and unimodal query formats. It further introduces hard negative samples, fine-grained reasoning categorization, and a reranking-rewriting enhancement strategy. Experiments across 23 state-of-the-art models reveal significant performance gaps in knowledge-intensive and reasoning-intensive tasks, with visual and spatial reasoning remaining key bottlenecks. The proposed enhancement strategies consistently yield measurable improvements.

benchmarkcomplex reasoninghard negatives

This work addresses the limitation of existing cross-modal retrieval models, which are predominantly confined to text–vision dual modalities and struggle to support composite queries involving text, vision, and audio. To this end, we propose OmniRet—the first efficient and high-fidelity tri-modal retrieval model capable of handling such composite queries. Our key innovations include an attention resampling mechanism that generates fixed-length compact representations for improved computational efficiency, and an attention-sliced Wasserstein pooling strategy designed to preserve fine-grained semantic information. We also introduce ACM, the first audio-centric multimodal benchmark. Built upon modality-specific encoders and a large language model architecture, OmniRet significantly outperforms current methods across 13 retrieval tasks and the MMEBv2 subset, demonstrating exceptional performance in composite querying, audio retrieval, and video retrieval, thereby validating its robust multimodal embedding capability.

audio-visual retrievalcomposed queriesmultimodal retrieval

Hot Scholars

ER

Elisa Ricci

University of Trento & Fondazione Bruno Kessler
Computer VisionDeep LearningRobotics
YD

Yuan Dong

Fudan University; Alibaba
Computer VisionMedical Image ComputingMachine Learning
MH

Ming-Hsuan Yang

University of California at Merced; Google DeepMind
Computer VisionMachine LearningArtificial Intelligence
XH

Xiaoshuai Hao

Beijing Academy of Artificial Intelligence,BAAI
vision and language