Composed Multi-modal Retrieval: A Survey of Approaches and Applications

📅 2025-03-03
📈 Citations: 0
Influential: 0
📄 PDF

career value

220K/year
🤖 AI Summary
To address the limitations in flexibility and accuracy of image/video retrieval amid the explosive growth of multimodal data, this paper presents a systematic survey of Compositional Multimodal Retrieval (CMR)—a paradigm that enables precise cross-modal search by composing reference visual content (images/videos) with textual modifications. We propose the first unified taxonomy for CMR and introduce a three-tier methodological framework encompassing supervised, zero-shot, and semi-supervised paradigms: supervised approaches emphasize data augmentation, architecture design, and loss optimization; zero-shot methods leverage external knowledge-guided modality translation. The framework integrates contrastive learning, modality alignment, prompt tuning, knowledge distillation, and multi-source synthesis, and is compatible with foundational models including ViT, CLIP, and BLIP. Evaluating over 100 studies, we demonstrate consistent improvements—12–28% higher retrieval accuracy—in applications such as product search, video understanding, and person re-identification, alongside significantly enhanced generalization compared to conventional cross-modal retrieval methods.

Technology Category

Application Category

📝 Abstract
With the rapid growth of multi-modal data from social media, short video platforms, and e-commerce, content-based retrieval has become essential for efficiently searching and utilizing heterogeneous information. Over time, retrieval techniques have evolved from Unimodal Retrieval (UR) to Cross-modal Retrieval (CR) and, more recently, to Composed Multi-modal Retrieval (CMR). CMR enables users to retrieve images or videos by integrating a reference visual input with textual modifications, enhancing search flexibility and precision. This paper provides a comprehensive review of CMR, covering its fundamental challenges, technical advancements, and categorization into supervised, zero-shot, and semi-supervised learning paradigms. We discuss key research directions, including data augmentation, model architecture, and loss optimization in supervised CMR, as well as transformation frameworks and external knowledge integration in zero-shot CMR. Additionally, we highlight the application potential of CMR in composed image retrieval, video retrieval, and person retrieval, which have significant implications for e-commerce, online search, and public security. Given its ability to refine and personalize search experiences, CMR is poised to become a pivotal technology in next-generation retrieval systems. A curated list of related works and resources is available at: https://github.com/kkzhang95/Awesome-Composed-Multi-modal-Retrieval
Problem

Research questions and friction points this paper is trying to address.

Efficiently search heterogeneous multi-modal data
Enhance search flexibility with composed multi-modal retrieval
Address challenges in supervised and zero-shot learning paradigms
Innovation

Methods, ideas, or system contributions that make the work stand out.

Integrates visual input with textual modifications
Categorizes into supervised, zero-shot, semi-supervised learning
Enhances search flexibility and precision