RidgeRank: Efficient Visual Document Reranking via Score Fusion and a Shallow Linear Readout

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high computational overhead of multimodal visual document reranking, where existing compression methods typically rely on annotated data or single-score paradigms that degrade accuracy. We propose an efficient, annotation-free reranking framework that jointly integrates retrieval and reranking scores via a closed-form optimal fusion rule to recover relevance signals, while introducing a shallow linear readout mechanism based on centered ridge regression to rectify intermediate representations. Evaluated across twelve datasets, the proposed method achieves NDCG@5 performance closely approximating that of full cross-encoders while accelerating inference by up to 48×. These results significantly expand the Pareto frontier between accuracy and latency for multimodal document reranking.
📝 Abstract
Multimodal language models rerank visual document retrieval results accurately, but scoring every candidate page at full cost makes them slow. Some methods that compress these rerankers need relevance labels to regain accuracy, and they rank by the reranker score alone. RidgeRank measures how much relevance signal the reranker score lacks and recovers it from the retriever score through a closed-form fusion rule. Maximizing a correlation objective gives the optimal fusion weight, along with the exact condition under which the reranker score by itself cannot reach that optimum. The reranker is further corrected by a single vector applied to an intermediate hidden state, obtained through one centered ridge regression onto the same model's full-depth scores on uncompressed pages. On 12 datasets drawn from ViDoRe 2 and ViDoRe 3, evaluated with two retrievers and two language model backbones, RidgeRank brings NDCG@5 to within 1.2 pp of a full cross encoder with speedups of up to 48 times, advancing the accuracy and latency Pareto frontier for visual document reranking.
Problem

Research questions and friction points this paper is trying to address.

visual document reranking
multimodal language models
score fusion
efficiency
compression
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Document Reranking
Score Fusion
Ridge Regression
Shallow Linear Readout
Multimodal Language Models
🔎 Similar Papers
No similar papers found.