🤖 AI Summary
This study addresses the prohibitive storage and comparison costs of multi-vector representations in visual document retrieval by proposing a parameter-free random soft token mechanism. During training, random unit vectors are appended to the encoder, and compact representations are learned via a late-interaction objective combined with a random sampling strategy. At inference, retrieval vectors are efficiently extracted without additional learnable parameters. Experiments demonstrate that, under an identical vector budget, this mechanism significantly outperforms existing readout strategies. Furthermore, substituting random inputs with zero tokens during inference preserves the performance gains, confirming that stochastic perturbations applied during training effectively enhance the quality of compact representations.
📝 Abstract
Visual document retrieval requires expressive representations to match queries with evidence distributed across text, tables, and page layouts. Multi-vector representations capture fine-grained information, but storing and comparing many vectors introduces substantial retrieval costs. In this paper, we introduce RandSlot, a simple approach to learn compact visual document representations with random soft tokens. During training, we append independently-sampled random unit vectors to query and document input sequences and resample them at every use, without introducing learnable soft-token parameters. The encoder contextualizes these auxiliary inputs with the original content to produce a small set of retrieval vectors. A standard late-interaction objective trains the encoder to extract relevant information under varying input conditions. Experiments with different backbone models show that RandSlot improves retrieval quality over alternative readout strategies under the same vector budget. Further analysis shows that these gains can persist when random soft tokens are replaced with zeros at inference, demonstrating that random inputs during training can improve compact retrieval representations even when inference no longer requires sampling.