🤖 AI Summary
This study addresses the limitation of existing unimodal metrics in accurately predicting the downstream performance of visual encoders within multimodal large language models, alongside experimental design and formulation flaws in prior cross-modal evaluations. Through extensive large-scale experiments, this work rectifies these shortcomings by proposing RAVEL, a training-free evaluation method. Built upon cross-modal nearest-neighbor retrieval, RAVEL enables efficient assessment of visual encoders without requiring additional training, thereby demonstrating the effectiveness of straightforward cross-modal metrics under rigorous experimental settings. Experimental results indicate that RAVEL achieves state-of-the-art performance across multiple benchmarks, substantially outperforming existing approaches and establishing a strong baseline for visual encoder evaluation.
📝 Abstract
Evaluating vision encoders requires metrics that reliably predict their downstream performance in multimodal large language models (MLLMs). Although recent studies have shown that cross-modal metrics can better capture such performance, unimodal metrics remain the dominant choice in practice. In this work, we revisit cross-modal evaluation of vision encoders through large-scale experiments. We identify important limitations in both the experimental design and methodological formulation of prior approaches. After addressing these limitations and introducing simple improvements, we propose RAVEL, a training-free method based on cross-modal nearest-neighbor retrieval. Despite its simplicity, RAVEL achieves state-of-the-art performance across our experiments, outperforming prior methods by a substantial margin. Our results demonstrate that simple cross-modal metrics, when evaluated under a careful and comprehensive setup, can provide a strong basis for evaluating vision encoders for MLLMs.