🤖 AI Summary
This study addresses the inherent trade-off between fine-grained representation and storage efficiency in multimodal embeddings by proposing ResComEmb, a general-purpose embedding framework. Methodologically, it introduces a novel residual homogeneity compression module to effectively eliminate visual token redundancy. By integrating dynamic resolution encoding with a length-adaptive bidirectional deferred interaction matching mechanism, the framework achieves efficient alignment for multimodal large language models. Experimental results demonstrate that ResComEmb comprehensively surpasses VLM2Vec-V2 on benchmarks such as MMEB. Notably, it outperforms ColQwen2.5 while utilizing only 37.5% of the token budget, significantly enhancing both retrieval effectiveness and computational efficiency in multimodal retrieval tasks.
📝 Abstract
Multimodal large language models (MLLMs) have shown strong potential for universal multimodal representation learning. However, existing methods either compress each input into a single vector, limiting fine-grained expressiveness, or retain long sequences of visual-token vectors, incurring substantial storage and interaction costs. To resolve this trade-off, we propose ResComEmb, a trainable framework for effective and efficient universal multi-vector multimodal embedding. ResComEmb first encodes each input at native dynamic resolution into ordered global, intermediate, and fine-grained views. After MLLM contextualization and embedding projection, a trainable Residual Homogeneity Compression (RHC) module reduces within-granularity redundancy and cross-granularity repetition under explicit visual token budgets. Then, ResComEmb introduces a length-adaptive Bidirectional Late-Interaction Matching mechanism for robust query-document scoring, which averages the strongest token-level matches in each direction and combines the two scores using a weight based on how many valid tokens each side has. Extensive experiments on MMEB, ViDoRe V1, and ViDoRe V2 show that ResComEmb produces higher-quality universal multimodal embeddings than VLM2Vec-V2, and outperforms ColQwen2.5 in visual document retrieval using only 37.5% of its full visual token budget, demonstrating a favorable effectiveness-efficiency trade-off.