Concepts Complement Dense Semantics: Learning Compact Sparse Spaces for Text-Image Retrieval

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of dense embedding spaces in obscuring fine-grained visual-textual information, which leads to insufficient cross-modal retrieval accuracy and a lack of explicit semantic grounding. To overcome this, we propose GRASP, a framework that leverages corpus-level concept mining to extract visual-textual concepts and construct a compact, conceptually sparse learning space. By integrating vision-language pre-trained models with a lightweight sparse prediction head, GRASP augments dense semantic matching with interpretable sparse evidence. This approach circumvents the reliance on redundant token spaces inherent in traditional methods. Extensive experiments demonstrate that GRASP surpasses state-of-the-art baselines in retrieval accuracy, achieving interpretable cross-modal retrieval that simultaneously delivers high precision, structural compactness, and explicit semantic grounding.
📝 Abstract
Cross-modal retrieval has been advanced by vision-language pre-trained models that encode images and texts into a shared dense embedding space. While dense representations effectively capture overall semantic similarity, they often obscure fine-grained visual-textual information needed for precise cross-modal matching. Recent methods introduce a learned sparse branch to complement dense matching with lexical evidence, but they rely on a redundant language-model token space and lack explicit grounding for sparse dimensions. We propose GRASP, a compact and grounded sparse learning framework that mines visual-textual concepts from the corpus. A lightweight sparse head is trained to predict concepts relevant to each image or text, yielding interpretable concept-level evidence that complements dense semantic matching. Extensive experiments show that GRASP improves retrieval accuracy over the state-of-the-art dense-sparse baselines while yielding a more compact and grounded sparse space.
Problem

Research questions and friction points this paper is trying to address.

cross-modal retrieval
dense-sparse representation
fine-grained matching
concept grounding
sparse space
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-modal retrieval
Sparse learning
Concept mining
Dense-sparse complementarity
Interpretable representations
🔎 Similar Papers
No similar papers found.
Y
Yoonseo Kim
Korea University, South Korea
J
Jungwoo Choi
Sungkyunkwan University, South Korea
C
Cheonyoung Park
KT Corporation, South Korea
Y
Youngwook Kim
KT Corporation, South Korea
Yongho Song
Yongho Song
Yonsei University
Conversational AIText RetrievalOpen-domain Question Answering
SeongKu Kang
SeongKu Kang
Korea University
Data miningText miningRecommender SystemInformation Retrieval