Are Benchmarks Reliable? Toward Structural Diagnosis via Sample-Level Capability Boundaries

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue that aggregated metrics in LLM benchmarks obscure sample-level reliability and introduce bias in model comparisons. To this end, it proposes BSDProbe, a framework that pioneers an instance-level capability boundary perspective. By analyzing repeated response trajectories to estimate model capability boundaries, combined with boundary location/width modeling and sequential consistency assessment, the framework diagnoses benchmark structures and selects highly discriminative subsets. The research reveals axis-conditional heterogeneity in benchmark reliability, overcoming the limitations of traditional leaderboard scoring. Experiments demonstrate significant stability variations across benchmarks such as GSM8K, with the selected compact subsets achieving up to an 8.58-fold improvement in model discriminability.
📝 Abstract
Evaluating large language models (LLMs) relies heavily on benchmark scores, yet aggregate metrics can obscure whether benchmark samples reliably support model comparison. We introduce \textbf{BSDProbe}, a sample-level framework for \emph{benchmark structural diagnosis} that estimates capability boundaries from repeated-response trajectories along ordered model axes. BSDProbe summarizes samples by boundary position, boundary width, boundary-signal validity, and order consistency, then aggregates them into benchmark-level structural profiles. Experiments on six benchmarks show that benchmark reliability is axis-conditioned and heterogeneous: GSM8K and MATH exhibit the most stable measurement structures, MMLU and TriviaQA are relatively stable but heterogeneous, while GPQA and PopQA show stronger axis-conditioned risks. These profiles remain consistent across Qwen3, Qwen2.5, and cross-model axes. BSDProbe further selects compact high-value subsets whose model discriminability reaches up to $8.58\times$ that of the full benchmark. These results suggest that reliable benchmark use requires examining sample-level capability boundaries beyond leaderboard scores.
Problem

Research questions and friction points this paper is trying to address.

benchmark reliability
large language models
model evaluation
sample-level diagnosis
capability boundaries
Innovation

Methods, ideas, or system contributions that make the work stand out.

Benchmark Structural Diagnosis
Capability Boundaries
Sample-Level Framework
Repeated-Response Trajectories
Model Discriminability
🔎 Similar Papers
H
Haiquan Hu
Beijing Normal University
Y
Yuzhu Liang
Beijing Normal University
W
Weicheng Tang
Beijing Normal University
Yanzeng Li
Yanzeng Li
Beijing Normal University
Y
Yao Shi
Beijing Normal University
Tian Wang
Tian Wang
Beijing Normal University
Edge ComputingInternet of ThingsSensor Cloud