Optimizing VLP-aligned Multimodal Intent Representation with Correct Visual Instantiation for Zero-Shot Composed Image Retrieval

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue in zero-shot composed image retrieval where reference noise and textual bias cause queries to deviate from the native space of vision-language pre-training (VLP) models. To this end, we propose VMIR-CVI, a unified framework that introduces Visual Multimodal Intent Reasoning (VMIR) integrated with few-shot chain-of-thought prompting to guide multimodal large language models (MLLMs) in generating aligned textual descriptions. Furthermore, it designs a Training-free Visual Instance Decoupling (TVID) module to eliminate noise interference and employs a lightweight Hybrid-modal Intent Alignment Fusion (HIAF) mechanism to optimize retrieval compatibility. Extensive experiments demonstrate that the proposed method significantly outperforms existing baselines on the CIRR, CIRCO, and FashionIQ benchmarks, achieving state-of-the-art performance.
📝 Abstract
ZS-CIR aims to retrieve a target image from a reference image and a modification text without paired supervision, typically by encoding composed queries as text-dominant representations within the image-text matching space of VLPs. However, queries reconstructed by visual pseudo-word learning or MLLM-based target reasoning often deviate from the native VLP representation space due to reference noise and coarse text fusion in the former, and verbose, weakly visually grounded descriptions in the latter. In this paper, we propose a unified ZS-CIR framework (named VMIR-CVI) to reconstruct multimodal composite queries from two complementary perspectives for optimizing VLP-compatible multimodal intent representation. First, it reasons and converts the multimodal intent into a unified textual description, aligning with the native text space of the VLP backbones to produce more retrieval-compatible textual queries. Second, it reconstructs the query representation with correctly decoupled visual instance cues, reducing reference noise while preserving target-relevant content. Specifically, a VLP-aligned Multimodal Intent Reasoning (VMIR) module injects few-shot VLP-style exemplars into chain-of-thought prompts, guiding the MLLM to generate target-consistent intent queries. A Training-free Visual Instance Disentanglement (TVID) module decouples fine-grained visual instances from global reference features without additional optimization. Finally, a lightweight Hybrid-modal Intent Alignment and Fusion (HIAF) module integrates the reasoned textual intent and disentangled visual cues into a unified hybrid-modal representation for robust ZS-CIR. Extensive experiments on three CIR benchmarks, namely CIRR, CIRCO and FashionIQ, show that VMIR-CVI significantly outperforms existing baselines and achieves new state-of-the-art performance. Code and trained models will be publicly released.
Problem

Research questions and friction points this paper is trying to address.

Zero-Shot Composed Image Retrieval
Vision-Language Pretraining
Multimodal Intent Representation
Representation Misalignment
Reference Noise
Innovation

Methods, ideas, or system contributions that make the work stand out.

Zero-Shot Composed Image Retrieval
Multimodal Intent Reasoning
Visual Instance Disentanglement
Vision-Language Pretraining
Chain-of-Thought Prompting
🔎 Similar Papers
No similar papers found.