Score
Designs, selects, and implements quantitative measures and procedures to assess the performance, accuracy, robustness, and fairness of models, algorithms, or systems; defines metric formulas, computes scores, evaluates statistical properties (e.g., variance, confidence intervals), and conducts comparisons and significance testing to interpret results and guide development decisions.
Traditional model evaluation relies on single-point metrics, failing to characterize performance stability and uncertainty. This paper proposes a small-sample (10–25 runs) uncertainty quantification framework tailored for high-reliability scenarios. It constructs empirical distributions of performance metrics via repeated stochastic experiments—encompassing random data splits, parameter initializations, and hyperparameter perturbations—and robustly estimates confidence intervals for metric quantiles using bias-corrected nonparametric bootstrap combined with quantile regression. To our knowledge, this is the first systematic approach enabling reliable confidence interval estimation for diverse metrics—including accuracy, F1-score, and MAE—in both classification and regression tasks under small-sample regimes. The method achieves high coverage (>90%) while maintaining narrow interval widths, thereby significantly improving robustness in model selection and enhancing decision-making credibility across multiple benchmark datasets.
In A/B testing, control variates and regression adjustment are widely used variance reduction techniques, yet their theoretical relationship remains unclear, their methodological frameworks are disjointed, and both have long been confined to design-driven paradigms. Method: This paper establishes, for the first time, a formal equivalence between these two approaches and proposes a novel grouped coefficient estimation method that unifies design-based and model-based estimation frameworks—enabling a paradigm shift from design-driven to model-driven inference. Contribution/Results: Theoretical analysis demonstrates improved estimation accuracy and statistical power. Empirical validation on millions of real-world experiments at ByteDance confirms efficacy: the proposed method has been fully deployed in its online experimentation platform, yielding an average 12.3% increase in statistical significance and a 19.6% improvement in detection sensitivity.
Data uncertainties—such as measurement errors, missing values, and erroneous links—undermine the credibility of policy decisions. Method: This paper proposes a decision-stability-oriented sensitivity analysis framework that shifts the analytical focus from parameter deviation to decision robustness. It introduces an interpretable, decision-level sensitivity metric and integrates counterfactual modeling, hypothesis-driven perturbation sampling, decision boundary tracking, and interactive visualization. Contribution/Results: Evaluated on two real-world policy domains—U.S. presidential vote prediction and childhood lead exposure assessment—the framework significantly enhances policymakers’ awareness of analytical robustness, explicitly delineates credible decision intervals, and provides an actionable confidence assessment tool for data-informed policymaking under data imperfections.
In quantitative pairwise comparisons, expert judgments are vulnerable to bribery-based manipulation, leading to distorted global rankings. Method: This paper formally defines the “targeted manipulation” problem for the first time and introduces a unified modeling framework integrating game theory and graph theory to characterize adversarial interventions. It proposes three polynomial-time solvable manipulation algorithms capable of precisely achieving desired rankings. Contribution/Results: Theoretical analysis demonstrates that even minimal bribery costs can significantly distort ranking outcomes. Furthermore, the study uncovers structural properties and inherent vulnerabilities of manipulation strategies, providing a theoretical foundation for detecting anomalous judgments and designing robust aggregation mechanisms. This work bridges a critical gap in the robustness literature on pairwise comparisons by establishing the first formal model of adversarial intervention.
Traditional binary correctness verification fails to capture quantitative system behaviors. Method: We propose the first automated toolkit for quantitative automata supporting six classical semantics—Inf, Sup, LimInf, LimSup, LimInfAvg, and LimSupAvg—and systematically address core decision problems: emptiness, inclusion, equivalence, and safety/liveness verification. Our approach introduces weighted transition modeling and a generalized value-function framework, integrating symbolic decision procedures, optimization solvers, and automata transformation techniques to enable extremal-value computation, safety-liveness decomposition, and real-time monitoring. Contribution/Results: Experiments demonstrate efficiency on inclusion checking, constant-function recognition, and online monitoring tasks. We release the first open-source benchmark suite for quantitative automata analysis, establishing a scalable, modular, and unified infrastructure for quantitative system verification.
This work addresses the limitations of existing formalisms for hyperproperties in capturing quantitative aspects inherent in real-world systems, such as numerical relationships in information flow control. To overcome this, the paper introduces Quantitative Hyper-Logic (QHL), a novel framework that reformulates hyperproperty specifications using measure theory, replacing classical Boolean quantifiers with measures to support nested quantitative structures. Leveraging Hoeffding’s inequality and extreme value theory, the authors develop an efficient statistical verification algorithm and provide rigorous analyses of sample complexity and statistical guarantees. Experimental evaluation on quantitative information-flow benchmarks demonstrates that QHL substantially outperforms conventional qualitative approaches, offering superior expressiveness and verification capabilities that better align with the demands of practical systems.
This work addresses the high cost of ground-truth evaluation in chemical and materials design, where existing machine learning surrogate models often lack reliability guarantees. Departing from conventional reliance on prediction accuracy metrics such as R²—which can paradoxically increase the risk of worst-case selections—the study proposes “rank preservation” as a core criterion for surrogate validation. It formally introduces the concept of “selection tax” and derives its theoretical upper and lower bounds. A safety certification framework for surrogates is established through selection-aware auditing, rank correlation analysis, and multi-task ground-truth validation. Experiments demonstrate that the proposed audit statistics achieve Spearman correlations of 0.80–0.99 with actual search performance, substantially outperforming R² (as low as 0.33). Certified screening strategies based on this framework reduce evaluation costs by up to 25-fold.
This study addresses the reliability of probabilistic uncertainty quantification (UQ) in software defect prediction, particularly its ability to reflect model performance and calibration—especially in cross-project settings, where systematic validation remains lacking. Through a large-scale empirical analysis of 16 classifiers across 36 within-project and 32 cross-project datasets, the work examines the relationships between five UQ metrics and six performance measures alongside three calibration metrics. It reveals, for the first time, a strong context dependency: within-project, UQ correlates strongly with false positive rate and AUC, but these correlations substantially weaken or even reverse in cross-project scenarios. Notably, high-performing models can still exhibit severe miscalibration. These findings indicate that UQ signals are not directly transferable and must be evaluated independently relative to specific objectives, using multidimensional calibration assessments.
Existing statistical model checking methods suffer from insufficient theoretical foundations and limited verification reliability. This work establishes the first comprehensive probabilistic-logical formal framework for the SCAN statistical model checker, integrating probabilistic model checking, statistical hypothesis testing, and formal verification techniques to rigorously characterize the property verification process of complex systems. By unifying these complementary approaches within a sound theoretical basis, the proposed framework not only addresses the foundational gaps previously present in SCAN but also significantly enhances its rigor and applicability. Consequently, it provides a robust guarantee for the reliability of SCAN when applied to the verification of real-world systems.
This study addresses the lack of a systematic framework for identifying critical input variables and conducting sensitivity analysis under uncertainty in complex simulations, particularly in military decision-making contexts. The authors propose a unified sensitivity analysis framework that integrates local and global methods—including variance-based, derivative-based, screening, and uncertainty quantification techniques—and strategically maps these approaches to specific decision objectives such as factor prioritization, fixing, variance reduction, and mapping. Innovatively, the framework introduces a “sensitivity audit” mechanism to enhance traceability of model assumptions and promote responsible model usage. By providing a structured guide for high-dimensional, complex simulation systems, this work significantly improves model interpretability, transparency, and the credibility of decisions derived from such models.