🤖 AI Summary
This study addresses the issue that aggregated metrics in LLM benchmarks obscure sample-level reliability and introduce bias in model comparisons. To this end, it proposes BSDProbe, a framework that pioneers an instance-level capability boundary perspective. By analyzing repeated response trajectories to estimate model capability boundaries, combined with boundary location/width modeling and sequential consistency assessment, the framework diagnoses benchmark structures and selects highly discriminative subsets. The research reveals axis-conditional heterogeneity in benchmark reliability, overcoming the limitations of traditional leaderboard scoring. Experiments demonstrate significant stability variations across benchmarks such as GSM8K, with the selected compact subsets achieving up to an 8.58-fold improvement in model discriminability.
📝 Abstract
Evaluating large language models (LLMs) relies heavily on benchmark scores, yet aggregate metrics can obscure whether benchmark samples reliably support model comparison. We introduce \textbf{BSDProbe}, a sample-level framework for \emph{benchmark structural diagnosis} that estimates capability boundaries from repeated-response trajectories along ordered model axes. BSDProbe summarizes samples by boundary position, boundary width, boundary-signal validity, and order consistency, then aggregates them into benchmark-level structural profiles. Experiments on six benchmarks show that benchmark reliability is axis-conditioned and heterogeneous: GSM8K and MATH exhibit the most stable measurement structures, MMLU and TriviaQA are relatively stable but heterogeneous, while GPQA and PopQA show stronger axis-conditioned risks. These profiles remain consistent across Qwen3, Qwen2.5, and cross-model axes. BSDProbe further selects compact high-value subsets whose model discriminability reaches up to $8.58\times$ that of the full benchmark. These results suggest that reliable benchmark use requires examining sample-level capability boundaries beyond leaderboard scores.