🤖 AI Summary
This work addresses the challenge of multimodal retrieval across text, images, and videos in enterprise search systems by proposing a unified late-interaction framework that operates without requiring modifications to existing backend architectures. The approach employs a multi-vector encoder to map heterogeneous modalities into a shared representation space and integrates a two-stage retrieval pipeline—comprising approximate nearest neighbor (ANN) candidate generation followed by accelerator-optimized exact MaxSim reranking—to enable efficient, fine-grained cross-modal search. As the first backend-agnostic multimodal solution deployed in a production system, it demonstrates strong empirical effectiveness on the ViDoRe V3 benchmark and achieves competitive ranking performance within a scalable Solr-based infrastructure.
📝 Abstract
We present AMES (Approximate Multimodal Enterprise Search), a unified multimodal late interaction retrieval architecture which is backend agnostic. AMES demonstrates that fine-grained multimodal late interaction retrieval can be deployed within a production grade enterprise search engine without architectural redesign. Text tokens, image patches, and video frames are embedded into a shared representation space using multi-vector encoders, enabling cross-modal retrieval without modality specific retrieval logic. AMES employs a two-stage pipeline: parallel token level ANN search with per document Top-M MaxSim approximation, followed by accelerator optimized Exact MaxSim re-ranking. Experiments on the ViDoRe V3 benchmark show that AMES achieves competitive ranking performance within a scalable, production ready Solr based system.