assess stability

Designs and executes procedures to evaluate whether estimated structures, model outputs, or analysis results remain consistent under perturbations or repeated execution. This includes building and running stability tests such as resampling (bootstrap, cross‑validation), pipeline reruns or repeat measurements, and pre‑specified evaluation criteria to quantify robustness and sensitivity.

assessstability

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.28
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$211K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the critical issue of model instability in software engineering optimization, which leads to substantial variability across repeated experiments and undermines both credibility and practical utility. Rather than treating instability as mere random noise, this work conceptualizes it as a quantifiable and manageable property that should be integrated into standard evaluation frameworks. By systematically modulating label usage, model complexity, and partition scoring strategies—combined with multi-objective optimization, causal intervention, data locality analysis, and model calibration—the proposed approach significantly enhances result consistency. Empirical evaluation demonstrates that the optimized configuration reduces the standard deviation of error by 22% on average and outperforms default settings in 119 out of 127 datasets, achieving a 4.8-fold improvement in result consistency.

model instabilitymulti-objective optimizationreproducibility

This study addresses the lack of systematic comparison between cross-validation and bootstrapping for assessing instability in clinical prediction models. Leveraging a cohort of 19,418 emergency department patients, it presents the first comprehensive evaluation of repeated five-fold cross-validation versus bootstrapping across varying events-per-variable (EPV) scenarios, using logistic regression and random forest models. Performance was assessed via AUC, calibration slope, large-scale calibration, and mean absolute prediction error (MAPE). Results indicate that when EPV ≥ 30, both methods yield comparable discriminative ability; however, cross-validation provides more accurate calibration estimates and significantly lower MAPE. These advantages render cross-validation particularly suitable for evaluating model instability across multiple algorithms, offering a dual benefit of internal validation and quantification of predictive stability.

bootstrapclinical prediction modelscross-validation

This study addresses the unreliable estimation of repeatability, between-laboratory, and reproducibility variance components under ISO 5725 standards when sample sizes are small or variance structures are extreme. To overcome this limitation, the authors propose a tailored Bootstrap resampling strategy adapted to a one-way random effects model. The approach refines point estimates by adjusting within-laboratory resampling and constructs confidence intervals via a two-stage resampling scheme integrated with bias-corrected and accelerated (BCa) techniques. Extensive simulations and validation using real data from ISO 5725-4 demonstrate that the proposed method substantially improves estimation accuracy and confidence interval coverage. It yields reliable, near-nominal or conservatively valid inferences for small- to moderate-sized experiments and clearly delineates optimal strategies across different practical scenarios.

bootstrapinterlaboratory precisionISO 5725

This work addresses the susceptibility of feature selection on small-to-medium scientific datasets to instability and optimistic bias induced by data leakage, which mutually exacerbate one another. The authors propose the first unified framework that integrates bootstrap-based stability selection with rigorous nested cross-validation, performing preprocessing and feature selection entirely within each inner fold to simultaneously yield highly stable feature subsets and unbiased performance estimates. Centered on stability as the primary optimization objective, the method supports binary classification, multiclass classification, and regression tasks, incorporates nine diverse algorithms, and includes a deterministic test suite to ensure reproducibility. Evaluated on three real-world datasets, the approach achieves predictive performance comparable to the best baselines while demonstrating significantly higher Jaccard stability than ANOVA F-test, recursive feature elimination, and Boruta.

data leakagefeature selection stabilitynested cross-validation

Latest Papers

What's happening recently
View more

Fixed-size benchmarking in model evaluation often fails to balance efficiency, statistical reliability, and diverse objectives, leading to either excessive resource consumption or unreliable results. This work proposes the first adaptive framework that integrates sequential testing into AI model evaluation, dynamically allocating evaluation data based on stopping criteria tailored for model ranking and selection tasks. By combining sequential hypothesis testing, minimum detectable effect analysis, and diminishing returns detection, the method achieves substantial gains in efficiency without compromising rigor. Empirical validation on the Open VLM Leaderboard demonstrates an 80% reduction in computational cost while maintaining a confidence interval width of 2.5 points, significantly enhancing both the practicality and scalability of model evaluation.

computational costfixed-size benchmarksmodel evaluation

This work addresses the high cost of ground-truth evaluation in chemical and materials design, where existing machine learning surrogate models often lack reliability guarantees. Departing from conventional reliance on prediction accuracy metrics such as R²—which can paradoxically increase the risk of worst-case selections—the study proposes “rank preservation” as a core criterion for surrogate validation. It formally introduces the concept of “selection tax” and derives its theoretical upper and lower bounds. A safety certification framework for surrogates is established through selection-aware auditing, rank correlation analysis, and multi-task ground-truth validation. Experiments demonstrate that the proposed audit statistics achieve Spearman correlations of 0.80–0.99 with actual search performance, substantially outperforming R² (as low as 0.33). Certified screening strategies based on this framework reduce evaluation costs by up to 25-fold.

experimental replacementmodel validationselection bias

Current perturbation-based construct validity audits are highly sensitive to implementation details, yielding conclusions that lack transparency and reliability. This work proposes a self-audit framework that systematically identifies and formalizes five classes of audit failure modes (F1–F5). A case study encompassing two open-source instruction-tuned models and five safety benchmarks reveals that none of the audited units satisfy confirmatory criteria, exposing systemic vulnerabilities in prevailing practices. To address this, the paper introduces a six-point due diligence gating mechanism that establishes actionable standards for disclosing and retaining high-assurance audit evidence, thereby substantially enhancing the credibility and reproducibility of auditing outcomes.

AI governanceaudit failurebenchmark validity

This work proposes a nonparametric method to assess the statistical significance of signal features—such as peaks and plateaus—in data and to detect multimodal structures in inter-event spacing distributions. The approach leverages run theory, employing a Markov chain recursion to precisely characterize the distribution of the longest runs. It integrates permutation testing with a bootstrap procedure tailored for continuous data, enabling a unified evaluation of both high- and low-intensity signal features. The key innovation lies in the first principled synthesis of run-length analysis, permutation tests, and continuous-data bootstrapping, which collectively facilitate accurate detection and localization of multimodal patterns without requiring parametric distributional assumptions, thereby effectively identifying salient morphological features in complex datasets.

bootstrap testmulti-modalitynon-parametric

Hot Scholars

XZ

Xing Zheng

Ph.D. of University of California, Riverside
Sensor fusionSLAMVIO
HW

Haobo Wang

Zhejiang University
Machine Learning
SY

Shenzhi Yang

Zhejiang University
machine learninglearning theorylarge language models
SX

Shixin Xu

Duke Kunshan Univeristy
machine learningmath biologyelectrodynamicsmoving contact lines