🤖 AI Summary
This study investigates whether multi-criteria evaluation feedback in adaptive development compromises the reliability of benchmarks for model selection and induces overfitting. By integrating statistical query theory with convex combination optimization, the authors analyze worst-case sample complexity and validate their findings on multi-task large language model benchmarks. The results reveal that the required sample size grows exponentially with the number of criteria, challenging the conventional assumption that benchmarks can be reliably reused. Furthermore, the study demonstrates that feedback from non-dominated tasks causes severe divergence between reuse-set and holdout-set scores, frequently yielding spurious winners. These findings provide a critical theoretical warning regarding the appropriate use of benchmarks in multi-criteria settings.
📝 Abstract
We study whether benchmarks reliably guide model selection as developers adapt to evaluation feedback across multiple criteria. We find that the worst-case test-set size needed to estimate the best score among $k$ adaptively chosen models, under any convex combination of the criteria, grows exponentially with the number of criteria, reaching the $\Theta(\sqrt{k})$ cost of answering $k$ adaptive statistical queries with only $O(\log k)$ criteria, at fixed accuracy and confidence. In attacks on multi-task large language model benchmarks with five to ten criteria, feedback restricted to nondominated task profiles produces large reused-to-held-out score gaps and frequent false winners. These results challenge a prominent explanation for prior observed reliable benchmark reuse---that developers mainly respond to convincing improvements over the current best---in rich-feedback settings, while leaving open how often ordinary model development encounters this vulnerability.