Score
Designs and runs systematic experiments and analysis pipelines that measure how changes in model choice, decoding and inference-time settings, and prompt variations affect model outputs, performance metrics, and error modes. Builds diagnostics, metrics, and visualizations to identify causes of performance variance (for example truncation, malformed outputs, or decoding hyperparameters) and to quantify sensitivity across prompts, models, and inference configurations.
Quantifying how input uncertainty propagates to model outputs remains a fundamental challenge in computational modeling. Method: This study systematically reviews and empirically compares prominent global and local sensitivity analysis (SA) techniques—including Sobol’, FAST, Morris screening, and local derivative-based methods—implemented via standard software packages, supporting both probabilistic modeling and distribution-free settings. Contribution/Results: We propose a practical decision framework that guides method selection based on problem characteristics, analytical objectives, and resource constraints—rejecting the notion of a universally “optimal” SA method and thereby addressing a critical gap in methodological implementation guidance. A reusable, open-source toolkit is developed to enhance the reliability and interpretability of uncertainty attribution. The framework and tools have been validated across multiple engineering and policy modeling applications, demonstrating robustness and scalability in real-world contexts.
When deploying machine learning models across heterogeneous environments, performance degradation often exhibits subgroup-specific heterogeneity—yet existing methods either explain only mean-level distributional shifts or isolate vulnerable subgroups without jointly identifying *where* degradation occurs and *why* it arises. This paper introduces SHIFT, the first hierarchical inference framework that unifies subgroup scanning, hierarchical causal inference, variable subset sensitivity analysis, and interpretable shift attribution. SHIFT simultaneously enables precise identification of degraded subgroups and disentanglement of underlying causes—distinguishing covariate shift from outcome shift. Evaluated on real-world deployments, SHIFT generates human-interpretable attributions of performance degradation and guides targeted interventions: it significantly improves performance for affected subgroups while avoiding negative transfer to others.
Fixed-size benchmarking in model evaluation often fails to balance efficiency, statistical reliability, and diverse objectives, leading to either excessive resource consumption or unreliable results. This work proposes the first adaptive framework that integrates sequential testing into AI model evaluation, dynamically allocating evaluation data based on stopping criteria tailored for model ranking and selection tasks. By combining sequential hypothesis testing, minimum detectable effect analysis, and diminishing returns detection, the method achieves substantial gains in efficiency without compromising rigor. Empirical validation on the Open VLM Leaderboard demonstrates an 80% reduction in computational cost while maintaining a confidence interval width of 2.5 points, significantly enhancing both the practicality and scalability of model evaluation.
Traditional model evaluation relies on single-point metrics, failing to characterize performance stability and uncertainty. This paper proposes a small-sample (10–25 runs) uncertainty quantification framework tailored for high-reliability scenarios. It constructs empirical distributions of performance metrics via repeated stochastic experiments—encompassing random data splits, parameter initializations, and hyperparameter perturbations—and robustly estimates confidence intervals for metric quantiles using bias-corrected nonparametric bootstrap combined with quantile regression. To our knowledge, this is the first systematic approach enabling reliable confidence interval estimation for diverse metrics—including accuracy, F1-score, and MAE—in both classification and regression tasks under small-sample regimes. The method achieves high coverage (>90%) while maintaining narrow interval widths, thereby significantly improving robustness in model selection and enhancing decision-making credibility across multiple benchmark datasets.
This work addresses the challenge that existing model evaluation methods often fail to reliably assess estimator quality in low-variance settings due to confounding between bias and variance or excessive sensitivity of statistical tests. To overcome this limitation, the authors propose a fault-tolerant evaluation framework that unifies bias and variance modeling through an adjustable tolerance parameter ε, enabling robust assessment of sample-efficient performance estimators within practically acceptable error margins. The framework integrates bias-variance analysis, fault-tolerant evaluation theory, and an adaptive ε-optimization algorithm, making it particularly well-suited for scenarios with low annotation costs. Experimental results demonstrate that the proposed approach provides a more comprehensive and reliable characterization of estimator behavior, significantly enhancing both the practical utility and stability of performance evaluation.
This study addresses hidden errors in large language model (LLM) evaluation arising from unquantified factors such as prompt rewrites, changes in judge models, or temperature variations, which can destabilize results and even reverse model rankings. The work presents the first systematic decomposition of these error sources, distinguishing between random variance—diminishing with increased data—and systematic bias sensitive to design choices. It proposes an optimized evaluation pipeline leveraging variance decomposition, few-shot estimation, and projection-based optimization. Empirical results across multitask benchmarks including MMLU demonstrate that, at equivalent computational cost, the method reduces estimation error by 50%, outperforms 73% of baseline evaluation protocols, and yields confidence intervals achieving near-nominal coverage—substantially enhancing evaluation robustness and mitigating noise overfitting.
This study investigates how the decodability of internal model representations dynamically evolves throughout pretraining and post-training, and whether erroneous decodability alone can reliably indicate discarded output information. Utilizing the Pythia model suite, the authors employ linear probing and steering intervention techniques to conduct cross-checkpoint comparative analyses of probe accuracy, steered responses, and error-correction mechanisms from early to late training stages. The work proposes an information-theoretic counterexample demonstrating that erroneous decodability is insufficient to establish the loss of output information. Furthermore, it reveals that while steering benefits improve progressively over the course of training, final-state decoders do not exhibit significant advantages. These findings offer novel perspectives for understanding the evolution of internal mechanisms within large language models.
本文提出一种精确测量大语言模型优化对输出质量影响的方法,通过校准的LLM评分系统,对比不同优化技术在相同提示下的表现。
研究通过12项任务和多种模型家族,揭示了输出格式对数据质量和模型能力评估的影响,并提出使用梯度特征等方法解决这一混淆问题。
为解决AI代理实验理解问题,引入WhatWorkedBench基准测试方法,通过预测组件更改后的结果准确性来评估,使用高斯过程提高效果恢复精度。
This study investigates whether tuning hyperparameters on test sets severely compromises the reliability of model evaluation and benchmark rankings. Through systematic experimental designs across multi-task benchmarks including MNIST, CIFAR, and GLUE, combined with statistical significance testing and ranking stability analysis, this work quantifies the actual impact of such practices. Challenging the conventional dogma that strictly prohibits test set tuning, the findings demonstrate that while this practice induces slight performance inflation, its magnitude frequently remains below the level of random noise and does not alter the relative ordering of models. By providing empirical evidence for re-examining this long-standing convention, this research advocates for a more open and transparent paradigm in evaluation reporting.