Hierarchical Compression of Vision-Language Model Benchmarks
This study addresses the prohibitive costs of benchmarking vision-language models (VLMs) and the underexplored nature of dataset compression by proposing PRIMEBench, a hierarchical compression framework. This work pioneers a visual-aware hierarchical architecture comprising a four-stage pipeline: data cleaning, category-representative selection, Visual-Aware Variance (VAW) pruning, and multimodal embedding analysis. By integrating visual dependency scores with model variance, the framework enables intelligent sample pruning. Experimental results demonstrate that PRIMEBench maintains high fidelity after removing over 97% of test samples, achieving optimal average ranking consistency with merely a 5% data retention rate. Consequently, this approach substantially reduces VLM evaluation costs while preserving benchmark reliability.