Learning Multimodal Embeddings with Evidence-Aligned Readout

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of effectively transforming semantic evidence generated by multimodal large language models into high-quality retrieval embeddings. To this end, it proposes EviAlign, a framework that couples semantic evidence generation with boundary-based readout within a shared model, constructing unified embedding vectors by aggregating unit boundary states. The research reveals a synergistic effect between semantic organization and readout strategies, demonstrating that consistent evidence structures significantly outperform length-position readout when employing boundary-based readout. Leveraging contrastive learning and joint training objectives, the proposed method achieves an average Recall@1 of 76.9 across 12 MMEB tasks using only 500K training pairs, delivering both superior performance and single-vector indexing efficiency.
📝 Abstract
Multimodal large language models can expose task-relevant evidence through generation, but producing useful evidence does not by itself determine how it enters a retrieval embedding. We study whether the semantic organization of that evidence can also specify where representations are read. To address this question, we introduce EviAlign, which couples Semantic Evidence Generation with Boundary Readout in a shared multimodal large language model. It organizes evidence into five semantic units, reads the contextualized state at each unit boundary, and aggregates these states into a single normalized embedding. Generation and contrastive retrieval objectives jointly train this shared structure. With the same trailing readout, semantic evidence and free-form CoT yield nearly identical retrieval performance, suggesting that evidence organization alone does not explain the full gain. A controlled $2\times3$ study compares consistent and permuted evidence organization across three readout strategies, using training targets with matched evidence spans. With five readout states and the same mean pooling, the advantage of consistent semantic organization grows from 0.65 points at length-based training positions to 2.39 at evidence boundaries, yielding a 1.74-point co-design interaction. Across 12 MMEB retrieval tasks, EviAlign achieves 76.9 average Recall@1 with 500K training pairs while retaining single-vector indexing and scoring.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Retrieval
Evidence Generation
Representation Readout
Multimodal Embeddings
Semantic Organization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Embeddings
Evidence-Aligned Readout
Semantic Evidence Generation
Contrastive Retrieval
Boundary Readout
🔎 Similar Papers
No similar papers found.
Zirong Chen
Zirong Chen
Vanderbilt University
cyber physical systemsnatural language processingartificial intelligencemachine learning
F
Fuda Ye
The Hong Kong University of Science and Technology (Guangzhou)
Enjun Du
Enjun Du
The Hong Kong University of Science and Technology (Guangzhou)
LLMData-MiningData-Centric AI
J
Junfu Pu
ARC Lab, Tencent
X
Xinlei Wang
Tencent Yuanbao
X
Xinyu Zuo
Tencent Yuanbao
L
Lisheng Duan
Tencent Yuanbao
H
Haijin Liang
Tencent Yuanbao
J
Jin Ma
Tencent Yuanbao
J
Jiachuan Wang
University of Tsukuba
Yongqi Zhang
Yongqi Zhang
Assistant Professor in HKUST(GZ)
Graph learningDrug discoveryDeep learning