🤖 AI Summary
This work addresses the high storage and memory costs of late interaction models, which generate numerous token-level vectors per document. To mitigate this, the authors propose a lightweight, pooling-aware fine-tuning approach that incorporates a compression objective during training, enabling flexible compression of multi-vector representations at inference time. By integrating k-means pooling with multi-factor training, the method demonstrates strong transferability across pooling strategies and datasets, and allows a single model to support multiple compression ratios. On the BEIR SciFact benchmark, the model maintains or even improves retrieval accuracy compared to an uncompressed baseline, despite achieving a compression rate of up to 83% (pooling factors 1–6).
📝 Abstract
Late interaction models have shown strong generalization capabilities, often outperforming much larger dense embedding models. One challenge to their widespread deployment is the large number of token vectors they produce per document and the associated storage and memory costs. Pooling tokens at inference time has shown great promise to reduce the vector count with limited effects on retrieval accuracy. Large-scale pooling-aware training has demonstrated even more impressive results at high compression rates. We propose lightweight fine-tuning as a practical alternative and find that even minimal pooling-aware training with k-means yields broad gains over inference-only pooling, shows evidence of transfer across pooling methods and datasets, and - with multi-factor training - produces a single model effective across different compression levels. Our strongest model outperforms the unpooled baseline on BEIR SciFact across pool factors 1-6, implying a vector compression rate of 83% at no cost to retrieval accuracy.