🤖 AI Summary
This study addresses the semantic mismatch between transformation-oriented language and target-state representations in zero-shot composed image retrieval (CIR). To this end, we propose ASAP-CIR, a training-free framework that reformulates the retrieval task as target-state reconstruction. Specifically, it leverages a frozen multimodal large language model to generate static target representations. Through atomic semantic weighting, calibrated evidence aggregation, and holistic state alignment, the method effectively suppresses interference from source images while preserving fine-grained visual constraints. Experimental results demonstrate that ASAP-CIR achieves superior performance on standard benchmarks such as FashionIQ and yields particularly significant improvements in multi-target scenarios on the CIRCO dataset.
📝 Abstract
Composed image retrieval (CIR) aims to retrieve a desired target image from a query consisting of a reference image and a modification text. This task exhibits an unusual representational asymmetry: the modification text specifies a transition from the reference state, whereas retrieval candidates depict completed target states. This creates a representation mismatch for zero-shot methods that query pretrained vision-language spaces directly with transformation-oriented language. We study this mismatch and reformulate zero-shot composed image retrieval as target-state reconstruction followed by retrieval. We instantiate this formulation with ASAP-CIR, a training-free framework that reconstructs a static target representation using a frozen multimodal large language model (MLLM). The representation combines multiple holistic descriptions with a variable set of importance-weighted atomic semantics, thereby preserving both overall target identity and fine-grained visual constraints. Retrieval then integrates holistic state alignment, atomic constraint grounding, and calibrated target-state evidence aggregation. A controlled text-only diagnostic shows that target-side static query formulations achieve more reliable retrieval than dynamic composed query formulations, particularly when source-state semantics must be suppressed or transformed. Experiments on FashionIQ, CIRR, and CIRCO further characterize the effectiveness and limitations of this representation principle, with the clearest gains on the multi-target CIRCO benchmark. These results show that how composed intent is represented before retrieval is a consequential design choice, distinct from the choice of retrieval backbone itself.