How Reusable Are Benchmarks with Richer Feedback?

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether multi-criteria evaluation feedback in adaptive development compromises the reliability of benchmarks for model selection and induces overfitting. By integrating statistical query theory with convex combination optimization, the authors analyze worst-case sample complexity and validate their findings on multi-task large language model benchmarks. The results reveal that the required sample size grows exponentially with the number of criteria, challenging the conventional assumption that benchmarks can be reliably reused. Furthermore, the study demonstrates that feedback from non-dominated tasks causes severe divergence between reuse-set and holdout-set scores, frequently yielding spurious winners. These findings provide a critical theoretical warning regarding the appropriate use of benchmarks in multi-criteria settings.
📝 Abstract
We study whether benchmarks reliably guide model selection as developers adapt to evaluation feedback across multiple criteria. We find that the worst-case test-set size needed to estimate the best score among $k$ adaptively chosen models, under any convex combination of the criteria, grows exponentially with the number of criteria, reaching the $\Theta(\sqrt{k})$ cost of answering $k$ adaptive statistical queries with only $O(\log k)$ criteria, at fixed accuracy and confidence. In attacks on multi-task large language model benchmarks with five to ten criteria, feedback restricted to nondominated task profiles produces large reused-to-held-out score gaps and frequent false winners. These results challenge a prominent explanation for prior observed reliable benchmark reuse---that developers mainly respond to convincing improvements over the current best---in rich-feedback settings, while leaving open how often ordinary model development encounters this vulnerability.
Problem

Research questions and friction points this paper is trying to address.

benchmark reuse
adaptive model selection
multi-criteria evaluation
rich feedback
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

adaptive model selection
multi-criteria benchmarks
test-set complexity
non-dominated feedback
benchmark reuse
🔎 Similar Papers
No similar papers found.