🤖 AI Summary
This study addresses the limitation that deterministic latent paths impose on embedding quality in multimodal retrieval. To overcome this constraint, we propose VaME, a framework that formulates latent reasoning as a learnable trajectory distribution. Specifically, VaME introduces variational latent reasoning to enable autoregressive exploration and optimizes stochastic trajectories by integrating reinforcement learning with semantic decoding rewards provided by a lightweight decoder, thereby transcending the limitations of conventional deterministic paths. Experimental results demonstrate that VaME significantly outperforms existing baselines on the MMEB-V2 benchmark, exhibits robust performance on MRMR tasks, and achieves at least a 4.25× improvement in inference speed.
📝 Abstract
Universal multimodal retrieval requires compact embeddings that preserve task-relevant semantic information across diverse modalities. Prior works have incorporated latent reasoning into multimodal embedding learning to refine this information before embedding extraction. However, most existing approaches remain confined to deterministic latent paths, without exploring alternative trajectories to discover better embeddings. Thus, we propose VaME (Variational Multimodal Embeddings), a framework that models latent reasoning as a learnable distribution over trajectories. Specifically, we first introduce Variational Latent Reasoning (VLR) to enable autoregressive exploration in latent space, guided by answer reconstruction through a lightweight decoder. Meanwhile, we augment the original embedding-token readout with a latent-fused embedding to facilitate exploration during subsequent reinforcement learning. Finally, we optimize latent reasoning over stochastic variational trajectories through reinforcement learning, using Semantic Decoding Reward (SDR) to favor semantically meaningful trajectories with interpretable decoded outcomes. On the 78-task MMEB-V2 benchmark, spanning image, video, and visual-document retrieval, VaME outperforms most explicit CoT-based models and all latent-reasoning baselines. VaME also demonstrates robust performance on reasoning-intensive benchmarks such as MRMR, with substantial gains after reinforcement learning. Importantly, VaME achieves these gains with at least a 4.25x inference speedup over the deterministic latent autoregressive baselines. The code will be made publicly available.