Score
Designs and evaluates methods that select or prune a small subset of benchmark datasets to estimate overall benchmark outcomes (such as model rankings or aggregate performance metrics) with reduced computation. Builds algorithms and estimators for subset selection/subsampling and analyzes subset-induced ranking quality and estimation error to preserve global model ordering while minimizing data used.
This work addresses the high cost of machine learning benchmarking by proposing a systematic framework to efficiently select small, representative subsets of datasets while preserving model ranking stability. The study presents the first comprehensive evaluation of various dataset selection strategies—including clustering, A/D-optimal experimental designs, random baselines, and a greedy farthest-first (FAFI) approach—on rank fidelity. It derives a theoretical upper bound on Spearman rank correlation error for FAFI and integrates bootstrap aggregation to yield statistically rigorous confidence intervals for comparing strategy performance. Empirical results demonstrate that as few as five datasets suffice to achieve 0.95 rank correlation in time series classification, significantly outperforming random selection in NLP tasks, though gains are limited in recommendation systems.
For generalized linear models (GLMs) under potential model misspecification in massive-data settings, existing subsampling methods suffer from low inferential efficiency and poor robustness. To address this, we propose a robust subsampling framework guided by prediction mean squared error (PMSE). Unlike conventional approaches assuming correct model specification, our method explicitly accommodates GLM misspecification by dynamically allocating sampling probabilities based on local PMSE estimates—thereby jointly accounting for data informativeness and model uncertainty. Theoretical analysis establishes consistency and asymptotic normality of the resulting estimator. Extensive simulations and real-world large-scale experiments demonstrate that our approach significantly outperforms state-of-the-art subsampling methods under model deviation, achieving a superior trade-off between computational efficiency and statistical accuracy. Consequently, it enhances both the reliability and practicality of statistical inference in large-scale, imperfectly specified modeling scenarios.
The NP-hard problem of best-subset selection in high-dimensional linear regression motivates this work. We propose an efficient, suboptimal algorithmic framework that integrates greedy search, regularized path tracking, and cross-validation-based model evaluation to yield a stable and scalable solution pipeline. Compared with mainstream heuristic approaches—including LASSO and orthogonal matching pursuit (OMP)—our method substantially reduces computational cost in ultra-high-dimensional settings (p ≥ 1000) while achieving superior trade-offs between model sparsity and predictive accuracy. Comprehensive benchmark experiments on both synthetic data and diverse real-world datasets demonstrate that our approach consistently attains higher solution quality and greater robustness than state-of-the-art baselines. By bridging efficiency and statistical reliability, the proposed framework establishes a new paradigm for high-dimensional sparse modeling.
For massive-scale multivariate linear regression where covariate dimensionality is low but the number of observations is extremely large, this paper proposes a subsampling design method based on the D-optimality criterion. Under mild assumptions on the covariate distribution, we integrate optimal experimental design theory with an equivalence theorem for constrained convex optimization to derive an analytically tractable acceptance–rejection rule; we further develop a computationally lightweight approximation algorithm. This work is the first to systematically embed optimal design theory into the subsampling framework for large-scale regression, substantially improving both statistical efficiency and computational speed. Simulation studies demonstrate superior performance over IBOSS. The resulting subsample-based estimator achieves asymptotic information-theoretic optimality, balancing estimation accuracy, robustness to model misspecification, and scalability to ultra-high-volume data.
Traditional static test sets inadequately evaluate foundation models’ diverse capabilities in open-ended scenarios. To address this, we propose ONEBench—a dynamic, extensible benchmarking paradigm that enables on-demand generation of customized evaluation suites targeting open capabilities, framing model assessment as a collective selection and aggregation process over sample-level tests. Our key contributions include: (1) the first unified, open-ended, and evolvable evaluation framework operating at the sample level; (2) a sparse measurement aggregation algorithm, a progressive sample pool construction mechanism, and a cross-modal unified interface (ONEBench-LLM/LMM); and (3) a robustness-aware scoring model with theoretical guarantees on identifiability and fast convergence. Experiments show that ONEBench achieves ranking stability >0.98 under 95% measurement sparsity, reduces evaluation cost by 20×, and attains >0.98 correlation with mean-score rankings on homogeneous data—enabling unified, efficient, and reliable assessment of both language and multimodal models.
This work addresses geometric data pruning methods that rely on neighborhood similarity assumptions, which inherently introduce selection bias. Discarding this assumption, we reformulate unbiased subset selection from first principles as a variance minimization problem. Through a linear programming perspective, we construct high-dimensional polytopes and derive closed-form pairwise variance expressions, enabling an efficient vertex-walking algorithm for label-agnostic data pruning with strictly guaranteed statistical unbiasedness. Experiments across multiple benchmarks demonstrate that the proposed method outperforms uniform sampling and mainstream geometric approaches in accuracy, exhibiting particularly superior performance under small selection budgets while effectively reducing stochastic gradient descent (SGD) variance.
This work addresses the high computational cost of evaluating large language models by reframing efficient benchmarking as a multivariate regression problem with feature selection. It introduces a novel approach that integrates minimum Redundancy Maximum Relevance (mRMR) feature selection with kernel ridge regression to identify a small, highly predictive subset of tasks from a full benchmark, enabling accurate estimation of overall model performance. Evaluated across multiple standard benchmarks, the method consistently outperforms existing strategies, achieving lower mean absolute error (MAE) and root mean squared error (RMSE), higher Spearman’s ρ and Kendall’s τ rank correlations, improved sampling efficiency, and greater consistency across random seeds—demonstrating a compelling combination of accuracy, robustness, and computational efficiency.
为降低软件工程代理回归测试成本,提出基于轨迹嵌入的子集选择方法,减少评估错误并节省计算资源。
Existing algorithm selection models exhibit limited generalization capabilities in real-world optimization scenarios, struggling to maintain consistent performance across diverse domains. This work presents the first systematic evaluation of cross-domain generalization between synthetic benchmarks (BBOB, CEC) and practical applications—specifically robotic trajectory optimization and UAV path planning—using an algorithm selection framework grounded in problem features and historical performance data, complemented by a carefully designed cross-benchmark experimental protocol. The study uncovers the failure mechanisms and success boundaries of current approaches when deployed in realistic settings, thereby providing crucial empirical insights for developing more robust and universally applicable algorithm selection systems.
Current evaluation practices for supervised learning models are often misleading due to an overreliance on single aggregate metrics, which neglect the alignment among data characteristics, task objectives, and real-world application contexts. This work reframes model evaluation as a context-dependent, decision-oriented process and systematically investigates—through controlled experiments—the impact of dataset properties, validation strategies, class imbalance, and asymmetric error costs on evaluation outcomes. Leveraging diverse benchmark datasets, multiple validation protocols, and multidimensional performance measures, the study uncovers common pitfalls such as the accuracy paradox, data leakage, and metric misuse. It proposes a structured evaluation framework explicitly aligned with operational goals, offering principled guidance for developing more robust, reliable, and trustworthy supervised learning systems.