inference sensitivity analysis

Designs and runs systematic experiments and analysis pipelines that measure how changes in model choice, decoding and inference-time settings, and prompt variations affect model outputs, performance metrics, and error modes. Builds diagnostics, metrics, and visualizations to identify causes of performance variance (for example truncation, malformed outputs, or decoding hyperparameters) and to quantify sensitivity across prompts, models, and inference configurations.

inferencesensitivityanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.39
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

"Who experiences large model decay and why?"A Hierarchical Framework for Diagnosing Heterogeneous Performance Drift

May 31, 2025
HS
Harvineet Singh
🏛️ University of California, San Francisco | Independent researcher

When deploying machine learning models across heterogeneous environments, performance degradation often exhibits subgroup-specific heterogeneity—yet existing methods either explain only mean-level distributional shifts or isolate vulnerable subgroups without jointly identifying *where* degradation occurs and *why* it arises. This paper introduces SHIFT, the first hierarchical inference framework that unifies subgroup scanning, hierarchical causal inference, variable subset sensitivity analysis, and interpretable shift attribution. SHIFT simultaneously enables precise identification of degraded subgroups and disentanglement of underlying causes—distinguishing covariate shift from outcome shift. Evaluated on real-world deployments, SHIFT generates human-interpretable attributions of performance degradation and guides targeted interventions: it significantly improves performance for affected subgroups while avoiding negative transfer to others.

Explains causes of decay via variable-specific shiftsIdentifies subgroups with significant performance decayProposes targeted actions to mitigate performance degradation

Fixed-size benchmarking in model evaluation often fails to balance efficiency, statistical reliability, and diverse objectives, leading to either excessive resource consumption or unreliable results. This work proposes the first adaptive framework that integrates sequential testing into AI model evaluation, dynamically allocating evaluation data based on stopping criteria tailored for model ranking and selection tasks. By combining sequential hypothesis testing, minimum detectable effect analysis, and diminishing returns detection, the method achieves substantial gains in efficiency without compromising rigor. Empirical validation on the Open VLM Leaderboard demonstrates an 80% reduction in computational cost while maintaining a confidence interval width of 2.5 points, significantly enhancing both the practicality and scalability of model evaluation.

computational costfixed-size benchmarksmodel evaluation

Traditional model evaluation relies on single-point metrics, failing to characterize performance stability and uncertainty. This paper proposes a small-sample (10–25 runs) uncertainty quantification framework tailored for high-reliability scenarios. It constructs empirical distributions of performance metrics via repeated stochastic experiments—encompassing random data splits, parameter initializations, and hyperparameter perturbations—and robustly estimates confidence intervals for metric quantiles using bias-corrected nonparametric bootstrap combined with quantile regression. To our knowledge, this is the first systematic approach enabling reliable confidence interval estimation for diverse metrics—including accuracy, F1-score, and MAE—in both classification and regression tasks under small-sample regimes. The method achieves high coverage (>90%) while maintaining narrow interval widths, thereby significantly improving robustness in model selection and enhancing decision-making credibility across multiple benchmark datasets.

Machine LearningModel EvaluationStability and Reliability

This work addresses the challenge that existing model evaluation methods often fail to reliably assess estimator quality in low-variance settings due to confounding between bias and variance or excessive sensitivity of statistical tests. To overcome this limitation, the authors propose a fault-tolerant evaluation framework that unifies bias and variance modeling through an adjustable tolerance parameter ε, enabling robust assessment of sample-efficient performance estimators within practically acceptable error margins. The framework integrates bias-variance analysis, fault-tolerant evaluation theory, and an adaptive ε-optimization algorithm, making it particularly well-suited for scenarios with low annotation costs. Experimental results demonstrate that the proposed approach provides a more comprehensive and reliable characterization of estimator behavior, significantly enhancing both the practical utility and stability of performance evaluation.

bias-variance tradeofffault-tolerant evaluationmodel performance estimation

This study addresses hidden errors in large language model (LLM) evaluation arising from unquantified factors such as prompt rewrites, changes in judge models, or temperature variations, which can destabilize results and even reverse model rankings. The work presents the first systematic decomposition of these error sources, distinguishing between random variance—diminishing with increased data—and systematic bias sensitive to design choices. It proposes an optimized evaluation pipeline leveraging variance decomposition, few-shot estimation, and projection-based optimization. Empirical results across multitask benchmarks including MMLU demonstrate that, at equivalent computational cost, the method reduces estimation error by 50%, outperforms 73% of baseline evaluation protocols, and yields confidence intervals achieving near-nominal coverage—substantially enhancing evaluation robustness and mitigating noise overfitting.

benchmark robustnessevaluation pipelinehidden uncertainty

Latest Papers

What's happening recently
View more

This study investigates how the decodability of internal model representations dynamically evolves throughout pretraining and post-training, and whether erroneous decodability alone can reliably indicate discarded output information. Utilizing the Pythia model suite, the authors employ linear probing and steering intervention techniques to conduct cross-checkpoint comparative analyses of probe accuracy, steered responses, and error-correction mechanisms from early to late training stages. The work proposes an information-theoretic counterexample demonstrating that erroneous decodability is insufficient to establish the loss of output information. Furthermore, it reveals that while steering benefits improve progressively over the course of training, final-state decoders do not exhibit significant advantages. These findings offer novel perspectives for understanding the evolution of internal mechanisms within large language models.

in-context decodinginformation-theoretic decodabilitymodel errors

This study investigates whether tuning hyperparameters on test sets severely compromises the reliability of model evaluation and benchmark rankings. Through systematic experimental designs across multi-task benchmarks including MNIST, CIFAR, and GLUE, combined with statistical significance testing and ranking stability analysis, this work quantifies the actual impact of such practices. Challenging the conventional dogma that strictly prohibits test set tuning, the findings demonstrate that while this practice induces slight performance inflation, its magnitude frequently remains below the level of random noise and does not alter the relative ordering of models. By providing empirical evidence for re-examining this long-standing convention, this research advocates for a more open and transparent paradigm in evaluation reporting.

benchmark integrityhyperparameter tuningmodel selection

Hot Scholars

JG

Jeremy Goldwasser

Statistics PhD Student, UC Berkeley
NLPInterpretabilityStatistical ML
IM

Ivana Malenica

Harvard University, U.C. Berkeley
StatisticsCausal InferenceMachine Learning
SD

Sandrine Dudoit

Division of Biostatistics and Department of Statistics, University of California, Berkeley
Applied StatisticsGenomics
PB

Peter B. Gilbert

Member, Fred Hutchinson Cancer Center, and Department of Biostatistics, University of Washington
Biostatisticsvaccine clinical trials