ResComEmb: Effective and Efficient Multimodal Embedding via Residual Homogeneity Compression

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent trade-off between fine-grained representation and storage efficiency in multimodal embeddings by proposing ResComEmb, a general-purpose embedding framework. Methodologically, it introduces a novel residual homogeneity compression module to effectively eliminate visual token redundancy. By integrating dynamic resolution encoding with a length-adaptive bidirectional deferred interaction matching mechanism, the framework achieves efficient alignment for multimodal large language models. Experimental results demonstrate that ResComEmb comprehensively surpasses VLM2Vec-V2 on benchmarks such as MMEB. Notably, it outperforms ColQwen2.5 while utilizing only 37.5% of the token budget, significantly enhancing both retrieval effectiveness and computational efficiency in multimodal retrieval tasks.
📝 Abstract
Multimodal large language models (MLLMs) have shown strong potential for universal multimodal representation learning. However, existing methods either compress each input into a single vector, limiting fine-grained expressiveness, or retain long sequences of visual-token vectors, incurring substantial storage and interaction costs. To resolve this trade-off, we propose ResComEmb, a trainable framework for effective and efficient universal multi-vector multimodal embedding. ResComEmb first encodes each input at native dynamic resolution into ordered global, intermediate, and fine-grained views. After MLLM contextualization and embedding projection, a trainable Residual Homogeneity Compression (RHC) module reduces within-granularity redundancy and cross-granularity repetition under explicit visual token budgets. Then, ResComEmb introduces a length-adaptive Bidirectional Late-Interaction Matching mechanism for robust query-document scoring, which averages the strongest token-level matches in each direction and combines the two scores using a weight based on how many valid tokens each side has. Extensive experiments on MMEB, ViDoRe V1, and ViDoRe V2 show that ResComEmb produces higher-quality universal multimodal embeddings than VLM2Vec-V2, and outperforms ColQwen2.5 in visual document retrieval using only 37.5% of its full visual token budget, demonstrating a favorable effectiveness-efficiency trade-off.
Problem

Research questions and friction points this paper is trying to address.

multimodal embedding
representation learning
visual token compression
fine-grained expressiveness
efficiency-effectiveness trade-off
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Embedding
Residual Homogeneity Compression
Late-Interaction Matching
Token Budget
Multimodal Large Language Models
🔎 Similar Papers
No similar papers found.
Z
Zijing Cai
University of Science and Technology of China
Y
Yuzhe Wang
University of Science and Technology of China
J
Jingxian Zhu
Hefei University of Technology
Fengbin Zhu
Fengbin Zhu
National University of Singapore
NLPIRLLMDocument AIAI + Finance
Richang Hong
Richang Hong
Hefei University of Technology
MultimediaPattern Recognition